Skip to content
roguelite labsAnthony Spezzano ↗
← Build log
Archive / earlier collection

AI-news notes.

The earlier site collected model, tool and industry news. This material is retained for review. Its claims and dates have not been reverified for this preview; follow the original sources before relying on an entry.

Page 3 of 7 · 123 archived entries

NVIDIA hits $4T / 2025 · Q3

NVIDIA closes at a $4 trillion market capitalization on July 10 — the first publicly traded company to reach that level, clearing Apple and Microsoft which had plateaued above $3T. The milestone is a direct product of the AI infrastructure buildout: every major lab and hyperscaler was deploying GPU clusters on a schedule measured in gigawatts, NVIDIA held the only production-grade training and inference hardware at the frontier, and the H100 had become the reserve currency of the AI economy. At $4T, NVIDIA was larger than the entire UK stock market and roughly equivalent to the GDP of Germany. The architecture moat — CUDA, NVLink, the NIM inference stack — had been built over two decades and had no realistic 18-month challenger. The market was not pricing near-term chip cycle revenue; it was pricing perceived multi-decade infrastructure dominance. Whether or not that view holds depends on how quickly AMD, custom silicon from Google and Amazon, and open alternatives like ROCM can close the software stack gap. In 2025, they hadn't. The $4T close was the clearest single data point that the AI compute race had become a permanent infrastructure category, not a temporary capex surge.

Original reference ↗
Grok 4 / 2025 · Q3

xAI ships Grok 4 and Grok 4 Heavy on July 9, unveiled via a livestream that drew 1.5 million concurrent viewers — the largest live audience for a model launch to that point. The benchmark that got the most attention was Humanity's Last Exam: Grok 4 Heavy scored above 50% on the text-only subset, the first model to clear that threshold on a benchmark explicitly designed to resist saturation. USAMO 2025 math proofs at 61.9%, ARC-AGI V2 at 15.9%. Trained on xAI's 200,000-GPU Colossus cluster using 6× more compute than Grok 3, with a 256K token context window. Native tool use is the architectural distinction: Grok 4 was trained to autonomously select its own web search queries and run a code interpreter mid-reasoning, rather than receiving tool calls as post-training injections. Grok 4 Heavy runs multiple reasoning agents in parallel at inference time, matching the test-time compute scaling pattern OpenAI had explored with o1-pro. The live audience and HLE score together mark xAI's transition from frontier competitor to challenger on the hardest tasks in the benchmark suite. The X data integration advantage — real-time access to the highest-velocity public information source without an API intermediary — remains the product moat no other lab can replicate.

Original reference ↗
Kimi K2 / 2025 · Q3

Moonshot AI releases Kimi K2 on July 11 — a 1.04 trillion parameter MoE with 32B active parameters per token, trained explicitly on long-horizon agentic tasks and tool use. Architecture: 384 experts (up from 256 in DeepSeek-V3), Multi-head Latent Attention, 128K context window. Modified MIT license — the most permissive license applied to a frontier-class parameter count at this date. Benchmarks at release: 65.8% SWE-bench Verified, 66.1 Tau2-Bench (agentic tool use), 76.5 ACEBench (agentic browsing). The Tau2 and ACEBench results matter more than SWE-bench here — K2 was designed for agent tasks, and it leads the open-weights category on those at release. The week of K2's launch: NVIDIA crosses $4T on July 10, Grok 4 ships July 9, and Kimi K2 ships July 11. Three simultaneous frontier events in 72 hours — the acceleration pace is no longer punctuated but continuous. The strategic signal is geographical: Moonshot AI was not a recognized open-weights player before this release. K2 arriving alongside DeepSeek V4 confirms that multiple Chinese labs have independently solved the training efficiency problem — and that the solution is replicating faster than any single-lab story suggested. Where DeepSeek pushed the open reasoning axis, K2 pushes the open agentic axis. Weights published on Hugging Face; same-day deployment on DeepInfra, Novita, Fireworks, and Baseten.

Original reference ↗
Windsurf: OpenAI deal collapsed, Google + Cognition split the pieces / 2025 · Q3

The Windsurf story in July 2025 is not a single acquisition — it's a three-way split that reshaped the agentic IDE landscape in one week. The sequence: OpenAI announced a ~$3B acquisition of Windsurf (Codeium) that was widely reported and expected to close; the deal collapsed under regulatory pressure and internal objections. With the deal dead, Google moved first: Google DeepMind hired Windsurf CEO Varun Mohan and his core leadership team and paid ~$2.4B for an intellectual property license — acquiring the technology and the people without buying the legal entity. What remained was a company with ~2 million developers, ~$82M annualized revenue, a VS Code fork with a capable agent (Cascade), and no CEO. Cognition — maker of Devin — stepped in and acquired that company for an undisclosed sum, likely a fraction of what OpenAI had offered. The deal gave Cognition a high-volume consumer coding product, Devin's enterprise positioning, and a combined developer base without any of the IP Google needed. OpenAI tried to buy a direct Claude Code competitor, failed, and watched Google absorb the talent and IP instead. The IDE landscape after the deal: Claude Code (Anthropic), Cursor (independent, $100M ARR), Cline (open-source), Amp (Sourcegraph), Codex CLI (OpenAI), and Windsurf/Devin (Cognition, Google holding the IP) as the five clear category competitors.

Original reference ↗
Gemini 2.5 Flash GA / 2025 · Q3

Google moves Gemini 2.5 Flash to general availability on June 17, completing the transition from experimental and preview releases. The model is the workhorse of the 2.5 family: 1M token context, built-in thinking mode, strong on reasoning and coding, priced at $0.30/$2.50 per million tokens — roughly a quarter of 2.5 Pro pricing for use cases where the additional capability headroom isn't required. Flash-Lite reached stable production status separately on July 22, at 1.5× the speed of Gemini 2.0 Flash at lower cost. The GA announcement reported a 25% improvement across internal benchmarks versus preview versions. The practical impact for developers was immediate: stable model IDs, committed deprecation timelines, and SLA-backed availability — the infrastructure properties that make a model usable in production rather than just interesting in a notebook. Google's two-track Flash/Pro strategy mirrors the competitive structure Anthropic established with Haiku/Sonnet/Opus, and means Google is now competing directly for the high-throughput API tier where most production tokens are actually spent, not just the prestige frontier tier where benchmarks are made.

Original reference ↗
Claude Sonnet 4.5 / 2025 · Q3

Anthropic's most capable agentic model at time of release. 61.4% on OSWorld — up from 42.2% on Sonnet 4 just four months prior. The jump in computer use performance without a price increase redefined what the mid-tier model tier meant. Strong reasoning, faster latency, the same $3/$15 pricing.

Original reference ↗
Gemini 2.5 Pro GA / 2025 · Q3

Google's Gemini 2.5 Pro moves from experimental to general availability, bringing its top-ranked long-context performance to production workloads. Consistently leads on coding benchmarks and remains the go-to for document-heavy and codebase-scale tasks requiring million-token context. Free tier via AI Studio.

Original reference ↗
The agentic IDE wave crests / 2025 · Q3

Cursor reaches $100M ARR — the fastest SaaS product to that milestone in history. Claude Code, Cline, Aider, Amp, and Codex CLI now each have distinct user bases and distinct use cases. The IDE has fractured: no single tool owns the workflow. Developers are running 3-4 agents per session, each assigned by task type.

Original reference ↗
Claude 4 — Opus 4 + Sonnet 4 / 2025 · Q2

Anthropic's Claude 4 generation ships May 22, 2025. Opus 4 benchmarks: 72.5% SWE-bench Verified, 43.2% Terminal-bench, 78.5% on GPQA Diamond, competitive with o3 and GPT-4.1 across the full evaluation suite — at the time of release, this placed it clearly above GPT-4o which sat at roughly 33% on SWE-bench. Sonnet 4 matches Opus 4 on SWE-bench at 72.7% at $3/$15 per million tokens versus Opus 4's $15/$75. Both models support a 200K context window. The key architectural unlock is extended thinking with tool use: prior Claude models had to reason first and then call tools, or vice versa; Claude 4 interleaves them mid-problem, so a reasoning trace can trigger a web search, incorporate the result, and keep reasoning — the difference is felt most on multi-step research and debugging tasks. Opus 4 is the first Anthropic model to trigger ASL-3 (AI Safety Level 3) protections from Anthropic's Responsible Scaling Policy — specifically, new CBRN-focused classifiers on inputs and outputs, plus enhanced security controls around weight storage, because Anthropic concluded it could no longer rule out the model providing meaningful uplift to someone attempting bioweapons work. Capable of autonomous sessions up to seven hours without human intervention.

Original reference ↗
Claude Code GA / 2025 · Q2

Launched as a limited research preview in February 2025, Claude Code went generally available on May 22, 2025 — the same day as Claude 4. The beta-to-GA delta was substantive: CLAUDE.md project memory (persistent instructions the agent reads on every session start), a hook system for triggering custom scripts before and after tool calls, native MCP integration for connecting external data sources and tools, context compaction for multi-hour sessions that would otherwise overflow the 200K window, and subagent support for parallel task execution. 'Unsupervised' in practice means the agent reads the full codebase, plans a sequence of file edits, runs tests, interprets failures, adjusts, and commits — without a human approving each step, only reviewing the final diff. The product hit $1B annualized run rate within six months of GA, and contributed directly to Anthropic's reported 4.5× revenue increase following the Claude 4 launch. 176 shipped updates in 2025 made it the most rapidly iterated product Anthropic had ever released.

Original reference ↗
Codex CLI / 2025 · Q2

OpenAI open-sources Codex CLI under Apache 2.0 in May 2025 — a terminal-native coding agent that accumulated 88K+ GitHub stars, making it one of the fastest-growing developer tools repos on the platform. The sandboxing model is OS-enforced: network access is off by default, file system access is scoped to the current workspace, and the approval policy requires explicit confirmation before any action outside those bounds. Three execution modes give callers control over autonomy: suggest (human approves every change), auto-edit (applies file edits without asking, asks before running commands), and full-auto (no interruptions until completion). The model layer runs the GPT-5-Codex family, optimized for repo-scale reasoning — by late 2025 this consolidated into GPT-5.2-Codex as the default. The key distinction from Claude Code: Codex CLI is intentionally lighter, with no built-in project memory file, no hook system, and a narrower agentic surface — it executes specific tasks cleanly rather than managing open-ended multi-hour sessions. Strong for scripting, sysadmin work, and precise surgical edits; Claude Code trades the lighter footprint for richer orchestration.

Original reference ↗
Tobi on AI at Shopify / 2025 · Q2

Lütke's internal memo goes wide: AI usage is now a baseline expectation at Shopify, not a differentiator. Performance reviews will account for it. The most direct statement from a tech CEO that AI fluency is a job requirement, not a bonus skill — and it came from a company with 10,000+ employees, not a startup.

Original reference ↗
Llama 4 — Scout + Maverick / 2025 · Q2

Meta ships Llama 4 on April 5 — the first Llama generation with MoE architecture and native multimodality via early fusion. Two models available for download: Scout (17B active / 109B total, 16 experts, 10M-token context via iRoPE) and Maverick (17B active / 400B total, 128 experts). A third, Behemoth (288B active / ~2T total), is previewed as the teacher model used for codistillation into Scout and Maverick; it's not released but already outperforms GPT-4.5, Claude Sonnet 3.7, and Gemini 2.0 Pro on MATH-500 and GPQA Diamond in preview testing. Scout fits on a single NVIDIA H100 with int4 quantization. Maverick runs on a DGX host (8× H100) and achieves ELO 1417 on LMArena's experimental chat mode — beating GPT-4o at half the active parameters. The 10M-token context in Scout, enabled by iRoPE (interleaved attention without positional embeddings), has no equivalent at open-weights scale at this date. Multimodality is early fusion: vision and language processed jointly from the input layer, not as separate towers bolted together. License: Llama 4 Community Agreement — custom, not fully open-source, but commercial use permitted. Weights at llama.com and Hugging Face. Meta's move to MoE is an implicit concession that dense scaling no longer works at the frontier: 400B total / 17B active is cheaper to train and serve than a 400B dense model. DeepSeek forced the recalculation; Meta's adoption confirms it.

Original reference ↗
Gemini 2.5 Pro Experimental / 2025 · Q2

Released March 2025 as an experimental preview — meaning rate-limited, no SLA, and subject to change without notice — Gemini 2.5 Pro debuted at #1 on the LMArena Chatbot leaderboard by the widest margin seen at that point. Benchmark profile: 84.0% on GPQA Diamond, 92.0% on AIME 2024 (pass@1), 86.7% on AIME 2025, 63.8% on SWE-bench Verified with a custom agent setup (Claude 3.7 Sonnet held the SWE-bench edge at 70.3%, but Gemini led on reasoning). Context window: 1 million tokens at experimental launch, with 2 million announced as forthcoming — no other model matched it at this price. What made it a genuine challenger rather than a narrow leader: it placed near the top on coding, math, long-context retrieval (94.5% on MRCR at 128K), and multimodal tasks simultaneously, where prior Google models had traded off against Claude or GPT-4o depending on the task type. The 'experimental' label carried a real cost in production — no uptime guarantees and aggressive rate limits — but developers used it anyway, which telegraphed how strong the demand signal was.

Original reference ↗
DeepSeek R1 / 2025 · Q1

DeepSeek releases the first open-source reasoning model trained via pure reinforcement learning to match OpenAI o1. 79.8% on AIME, 97.3% on MATH-500. Permissive license, full 671B weights published. Shattered the assumption that reasoning capability required closed training data or RLHF at proprietary scale. The week it shipped, every frontier lab's stock dropped.

Original reference ↗
Vibe coding / 2025 · Q1

Karpathy coins the term in a tweet: describe a project in natural language, accept AI-generated code without reading it, iterate on behavior not syntax. Within weeks, Merriam-Webster adds it to the dictionary. Collins names it Word of the Year 2025. A name for a practice millions of people were already doing — which made them realize they were allowed to do it.

Original reference ↗
Claude 3.7 Sonnet / 2025 · Q1

Anthropic's first hybrid reasoning model: standard fast responses for most tasks, extended thinking mode for hard problems. The first Anthropic model to show o1-style chain-of-thought, interleaved with tool use. On SWE-bench Verified: 62.3%. Marked the transition from Claude as a chat model to Claude as a reasoning system.

Original reference ↗
Grok 3 / 2025 · Q1

xAI ships Grok 3, trained on 10x the compute of Grok 2. Tops major reasoning benchmarks at launch. Competitive on math and graduate-level science. Access to real-time X data gives it a live information advantage no other frontier model has. The first xAI model that felt like genuine frontier competition.

Original reference ↗
OpenAI Agents SDK / 2025 · Q1

OpenAI open-sources a multi-agent orchestration framework with first-class primitives for handoffs, guardrails, and tool use across agent graphs. The Python SDK is lean enough to actually use. Arrives alongside Responses API, which replaces the older Completions and Assistants endpoints. The clearest signal that OpenAI sees agent orchestration — not raw models — as the product layer.

Original reference ↗
Computer use / 2024 · Q4

Anthropic ships Claude 3.5 computer use in public beta: Claude can see your screen, move the cursor, click, type, and navigate applications like a human operator. First frontier model to offer desktop automation as a first-class capability. Simultaneously, Claude 3.5 Haiku ships — fast, cheap, and surprising: 40.6% on SWE-bench, beating models twice its size.

Original reference ↗