Executive Summary
Today's sources share one lesson: a component's efficiency claim tells you little until you place it in the system around it. SemiAnalysis shows that sparse attention does not reduce memory capacity by itself. It also shows that GPU cost-per-token rankings reverse depending on the latency SLA. Willison documents an agent that could send messages but could not verify anything, and its unverified claim caused measurable reputational damage. Jones describes a narrow classifier that removes output-token cost entirely for purely decisional work. For enterprise buyers, the practical result is a decomposition discipline. Split workloads by latency tolerance, by whether they need generation or only a decision, and by whether the agent's claims can be verified. Then buy and govern each class separately. Anyone who evaluates a model, chip, or agent in isolation is currently being misled by its headline numbers.
What Changed
- Sparse attention savings depend on the serving stack. Top-k selection still needs the full KV cache held in HBM. The capacity relief comes from serving-engine tiering (SGLang HiSparse offloading to host DRAM), not from the attention mechanism. As concurrency goes from 8 to 16, GPU-resident cache reuse falls from 90.3% to 54.8% and host-memory reuse rises from 6.0% to 40.3%, while overall hit rate stays above 95%.
- No GPU wins across the whole range. GB200 beats MI355X/ATOM by 5 to 12% at high throughput. ATOM wins under tight TTFT, where GB200's fastest configs wait 14 to 19 seconds for the first token against ATOM's roughly 1.1 seconds. GB300 is 26% cheaper only if you accept a 10 second TTFT.
- Serving engineering now moves results as much as silicon. CUDA graph consistency cut TPOT from 40ms to 22ms. Chunked prefill pipelining nearly doubled throughput. TileRT's single-kernel decode gave MI355X a 40% interactivity edge over GB300 NVL72.
- Chinese labs are co-designing models for non-Nvidia silicon. GLM-5 uses half DeepSeek's query-head count, and Z.ai had day-0 Moore Threads support. Both point to deliberate optimization for the MTT S4000.
- Decisional inference is becoming its own product category. Jev returns a discrete choice with zero output tokens. It is built to plug into existing agent stacks as a subroutine.
Cross-Expert Synthesis
SemiAnalysis and Jones reach the same conclusion from opposite directions. Workload shape, not model capability, should drive infrastructure spend. SemiAnalysis splits by latency tolerance: interactive work versus batch or background agentic work. Jones splits by output modality: generative versus decisional. Put together, they give a four-way grid, and each cell has a different cheapest option. That can be a different chip, a different serving configuration, or a different model class. A single enterprise "LLM endpoint" is overpaying in at least three of the four cells.
Willison's incident is where this optimism runs out. The Muse agent's failure was a decision ("am I able to confirm presence?") presented as generation ("Yep I'm here!"). Cheap classifiers make it tempting to add more autonomous gates. But a classifier cannot verify physical-world state any better than an LLM can, because verification needs a sensor, a calendar, or a system of record. Making decisions cheaper increases the number of automated assertions, and that raises the cost of ungrounded ones. Cheap decisions only reduce risk when the input they read is actually grounded.
Enterprise Implications
Procurement based on cost-per-token without a stated TTFT SLA is now demonstrably wrong. The same benchmark supports opposite vendor choices at 2 seconds and at 10 seconds. RFPs that request "tokens per dollar" without latency bands will produce misleading answers from vendors.
Memory sizing has to include host DRAM. Long-context agentic serving at concurrency pushes almost half of cache reuse into host memory. Nodes specified with minimal DRAM and weak CPU to GPU bandwidth will fail to deliver the efficiency that sparse attention models advertise.
Model efficiency claims, including "sparse," "efficient," and "long-context cheap," cannot be used for capacity planning unless a specific serving engine and version is named. Serving-stack choice is now an architecture decision, not an operations detail.
Hardware concentration risk is easing on the supply side. Chinese labs targeting Moore Threads, and AMD winning latency-constrained segments, both reduce the argument for single-vendor lock-in. That is useful leverage in Nvidia negotiations even for buyers who never deploy AMD.
Agent governance has a gap between communication authority and verification authority. Any agent that can message customers, suppliers, or counterparties on someone's behalf can also, as Willison shows, issue apologies and remediation. Those are commitments that nobody approved.
What To Do About It
- Build a workload grid before any hardware conversation. Classify each workload as interactive or batch, and as generative or decisional. Price each cell separately. Route batch agentic work to throughput-optimized configs, and consider GB300 where 10 second TTFT is acceptable.
- Audit LLM calls for decisional work. Routing, triage, approve/reject, and parameter selection are candidates to move to narrow classifiers. Pilot Jev or equivalents on one high-volume gate and measure accuracy against the current LLM before moving anything else.
- Require the serving stack in every efficiency claim. Name the engine, version, offload configuration, and concurrency level, or treat the number as marketing.
- Set a pre-deployment gate for customer-facing agents: every assertion of real-world state must trace to a verifiable source, or be phrased explicitly as uncertain. Also require explicit authorization for agent-issued apologies, commitments, and remediation.
- Customer question worth asking: "What TTFT does this workload actually need, and who signed off on that number?" Most teams cannot answer it, which means their GPU cost comparison is unfounded.
Risks and Blind Spots
The InferenceX figures are a single dated snapshot. Serving-engine updates regularly shift results by 40 to 98%, so today's rankings should not be written into multi-year contracts. Jev's adoption superlative is self-reported and only days old, and its quoted pricing is ambiguous in the source. Treat the cost case as a hypothesis to test in a pilot, not a baseline. The Moore Threads inference is SemiAnalysis's reading of architectural evidence, not a confirmed Z.ai statement. The Muse incident is one anecdote, but the failure mode is structural, and a single-case origin is not a reason to dismiss it. Last, the workload-decomposition discipline adds routing complexity and more models to operate. For smaller enterprises, the operational overhead may cancel the per-token savings, so measure total cost of ownership before breaking up a working monolith.