Executive Summary
Four sources, one underlying variable: verifiability. Where AI operates against clean, checkable feedback signals (chip design, weather physics, code compilation, legal citation-checking) capability is compounding fast and enterprises are already shipping production systems on it. Where it doesn't (organizational politics, most SMB workflows, general judgment) adoption is stalling into "glorified chatbot" territory regardless of model quality. The token-economics data confirms the same split from a different angle: usage volume is migrating to cheap open-weights models, but spend is not, because frontier labs still own the hardest, highest-value, least-verifiable tasks. The strategic read for an infrastructure architect is not "AI is getting more general," it's "AI capability is stratifying by task-verifiability, and the winners are the organizations that build the management and integration layer around that stratification before their competitors do."
What Changed
- Open-weights models (DeepSeek, Qwen, Kimi, GLM) now exceed closed models in raw token share on Vercel's platform telemetry, with DeepSeek (25.2%) surpassing Anthropic (24.5%) — but Anthropic still captures 64.6% of spend versus DeepSeek's 2.8%, a 23x gap. Volume and value have decoupled.
- Agent token consumption is up 14x from February to August 2026 and now runs 5x human token volume. This is a straight rebuttal to the "agents replace headcount" purchasing thesis still driving most enterprise budgets.
- Legal-sector Codex usage is up 108x since January, the sharpest vertical adoption curve reported, because legal work is unusually verifiable (citations either check out or don't).
- Anandkumar's Fourier Neural Operator weather model, trained on ~50K samples, is in live production at ECMWF, correctly called Hurricane Lee's landfall ahead of standard forecasts, and has become the architectural basis for a major climate model — a concrete case of AI replacing a decades-old, supercomputer-class scientific pipeline, not just augmenting it.
- A production incident (Pocket OS: an over-scoped agent credential deleted a live database in 9 seconds, 30 hours to recover) is now the reference failure mode for unmanaged agent permissions, playing the same role the early ransomware incidents played for security budgets.
Cross-Expert Synthesis
Every source is independently describing the same fault line from a different altitude. Greenblatt's argument is that transformative capability only requires mastery of verifiable, feedback-rich domains, not general judgment. Anandkumar's neural operators are the mechanistic proof of exactly that claim inside physical science: strong priors (conservation laws, geometry) plus a verifiable loss function let a model with three orders of magnitude less data than a transformer would need outperform brute-force scaling. Jones's adoption data is the labor-market version of the same pattern: agents scale fastest in legal and code because correctness is checkable, and stall everywhere else. Berman's pricing data is the market's version: frontier labs charge a 23x spend premium precisely for the hard, ambiguous, low-verifiability tasks that commodity open-weights models can't reliably do, while commodity models absorb the high-volume, low-difficulty, checkable work.
The tension worth naming: Anandkumar argues AI-for-science should get lighter regulatory treatment than agentic/language-model AI, precisely because it's verifiable and physics-constrained. Greenblatt's warning cuts the other way — the domains where AI is most capable (hardware R&D, engineering, fab-scale buildout) are exactly where humans are most likely to lose the ability to audit what's being built, because volume and opacity scale together even when correctness is locally verifiable. Local verifiability (does this simulation match physics) does not imply global auditability (do we understand what the aggregate R&D output is doing to the industrial base). BlueAlly customers building on either neural-operator-style scientific AI or agentic R&D pipelines will hit this gap before regulators catch up to it.
Where AI Is Heading
Two parallel tracks, not one. Track one: agentic AI keeps expanding total work volume (Jevons effect, per Jones) inside enterprises that can afford to build a management layer around it — access control, permissioning, oversight roles. Track two: physical-world and hardware AI (Greenblatt's chip design and fabs, Anandkumar's neural operators) advances on small, high-quality, physics-structured data rather than internet-scale data, which changes who can compete. An energy company or materials firm sitting on 10,000 well-instrumented simulation runs is now a credible AI competitor to a hyperscaler, because the bottleneck shifted from data volume to data structure. Meanwhile the open-weights/closed-weights market is settling into three durable tiers (commodity, enterprise-fine-tuned specialist, closed frontier generalist) rather than converging to one winner — architecture decisions should assume this tiering persists for years, not treat it as a transitional state before consolidation.
What Enterprise Customers Should Care About
- Their AI ROI model is probably wrong if it assumes fewer agents means fewer people. The realistic outcome is the same headcount doing more, plus a new oversight function nobody budgeted for.
- Cost-per-token comparisons between open and closed models are misleading; cost-per-completed-task is the only number that matters, and cheaper models often need more tokens to reach the same answer.
- Sending proprietary operational data through a hosted frontier model is a competitive exposure decision, not just a procurement one — the lab gets visibility into how the business runs.
- If they have structured physical or sensor data (manufacturing, energy, chemicals, logistics), they may already have enough data to deploy physics-informed models in production, years earlier than they assume, without waiting for more data collection.
What BlueAlly Should Say
Lead with the verifiability framework, not with "AI is transformative." Customers have heard the hype; they haven't heard someone tell them precisely which of their workflows are structurally suited to AI automation right now (checkable, feedback-rich) versus which will keep disappointing them regardless of model choice (judgment-heavy, ambiguous). Pair that with a blunt correction to the headcount narrative: budget for the management layer up front, or the Pocket OS failure mode becomes their failure mode. On infrastructure, be the vendor that talks candidly about the platform-risk and chip-dependency tradeoffs in open-weights adoption instead of pitching a single-vendor stack — that candor is a differentiator against system integrators still selling one-cloud, one-model narratives.
Infrastructure Implications
- Agent oversight requires new infrastructure: permissioning, audit logging, kill switches, and staged rollout gates, sized for the 10-20 agents per 10-person team density Jones projects by 2027, not for today's pilot-scale deployments.
- Fine-tuning open-weights models in-house (the Harvey/Airbnb/Perplexity/Cursor pattern) requires a serious self-hosting and fine-tuning capability that most enterprises don't have and will need to buy rather than build from scratch.
- Physical/industrial AI workloads (neural operators) run tens of thousands of times faster per task than legacy simulation, which inverts the usual compute-scaling assumption: expect lower per-task compute intensity but continued or growing aggregate demand, since Anandkumar states compute remains her bottleneck even after the efficiency gain.
- Chinese-chip-co-designed open-weights models introduce a supply chain dependency that infrastructure architects need to track independently of the AI-model-selection decision — it's a hardware procurement risk wearing a software vendor's clothes.
Security and Governance Implications
- Unmanaged agent permissions are now a demonstrated production-outage vector, not a theoretical risk (Pocket OS). Credential scoping and blast-radius limits for agent tooling should be treated with the same rigor as service-account permissions in any other automated system.
- Platform risk from closed-model vendors — the lab's visibility into a customer's proprietary operational data — belongs in vendor risk assessments and contract negotiations, not just cost comparisons.
- The interpretability gap Greenblatt describes (large volumes of AI-conducted R&D that humans can't fully audit) is an operational governance issue for any enterprise running AI-driven engineering or R&D pipelines at scale, independent of any AI-safety policy debate.
- Anandkumar's push for separate, lighter regulatory treatment of AI-for-science is a live policy fight (she has UN advisory involvement); enterprises deploying physics-informed models should track this rather than assume today's AI governance frameworks (built around language models and agents) will apply to them unchanged.
Sales Talk Tracks
- "Your agent ROI case is probably built on a headcount assumption that the token data doesn't support. Let's model what the oversight layer actually costs, so the business case survives contact with reality."
- "We'll tell you which of your workflows are verifiable enough for agentic automation to actually work today, and which ones will burn budget for another 18 months no matter which model you pick."
- "If you're evaluating open-weights fine-tuning for cost or data-sovereignty reasons, we'll also flag the platform-risk and supply-chain tradeoffs the vendor pitch won't mention."
- "If you're sitting on years of sensor, simulation, or process data, you may already have what you need to deploy physics-informed AI in production, not in five years."
Customer Discovery Questions
- Which of your current or planned agent deployments have a checkable, verifiable success criterion, and which are being deployed on judgment calls that nobody can objectively grade?
- What's your actual plan for agent credential scoping, audit logging, and permission review as agent count scales past the pilot stage?
- Has anyone modeled what data leaves your environment when you route proprietary workflows through a hosted frontier model, and what that vendor could do with it later?
- Do you have structured physical, sensor, or simulation data sitting unused that could support a physics-informed model instead of a data-hungry general-purpose one?
- Who in your organization owns the "agent management tax," or is that decision still undefined?
Potential BlueAlly Service Opportunities
- Agent governance-as-a-service: permissioning, logging, and incident-response tooling built specifically for agent fleets, positioned against the Pocket OS failure mode.
- Open-weights fine-tuning and self-hosting engagements for customers chasing the Harvey/Cursor pattern, including a platform-risk and chip-dependency assessment as part of the engagement.
- Task-verifiability audits: a structured assessment of which customer workflows are automation-ready today versus not, sold as a prerequisite to any agent deployment engagement.
- Physics-informed model pilots (neural-operator-class) for customers in energy, materials, chemicals, or manufacturing who have underused simulation or sensor archives.
Risks and Blind Spots
- BlueAlly's own sales motion may still be implicitly selling the headcount-reduction story Jones's data contradicts; that pitch will erode trust with technically literate buyers who've seen the OpenRouter numbers.
- The three-tier model market means there is no single "right" model recommendation to standardize on; a one-size-fits-all vendor recommendation will age poorly within a year.
- Nobody in this source set has a good answer for the interpretability/audit gap in AI-driven R&D; treating it as solved or someone else's problem is itself a blind spot worth flagging to customers rather than papering over.
Contrarian Viewpoints
Berman himself reverses his prior threat model mid-analysis: he now judges Chinese-chip dependency in open-weights infrastructure as a more serious risk than the AI-power-concentration narrative (OpenAI/Anthropic dominance) he previously worried about. That reversal is worth surfacing to customers directly rather than smoothing over — the geopolitical risk profile of "cheap open models" is not static and the current framing may itself be outdated within a year.
Separately, Anandkumar's push to regulate AI-for-science more loosely than agentic AI sits in direct tension with Greenblatt's warning that R&D-heavy AI is precisely where human audit capacity is most at risk of being outpaced. Both can't be fully right: if verifiable-domain AI is where transformative, hard-to-audit industrial buildout happens fastest, that argues for more oversight of science/engineering AI, not a lighter regulatory touch. Enterprises should not assume the "safe because verifiable" framing settles the governance question.