SRAM inference racks, a $110 billion cheque, and search as a gateway tool
Nvidia put Groq 3 LPX into production for ultra-fast agent tokens, OpenAI’s record $110 billion round recast who funds the frontier, and Cloudflare turned web search into a first-class AI Gateway API. Speed, capital and grounding — the three constraints of production agents — all moved.
Key takeaways
- Groq 3 LPX is a production, rack-scale decode engine beside Vera Rubin, not a science project; early cloud access is through specialists such as Nebius.
- OpenAI’s $110 billion round at a $730 billion pre-money valuation ties the company more tightly to Amazon and Nvidia capacity.
- Cloudflare’s Web Search API lets teams buy grounded search with the same billing and logs as their other model calls.
Nvidia Groq 3 LPX enters full production for low-latency agent inference
On August 24 Nvidia said Groq 3 LPX, a rack-scale SRAM inference accelerator for the Vera Rubin platform, is in full production, citing 3,400 output tokens per second on Gemma 4 31B at 100,000-token context in Artificial Analysis numbers.
At Hot Chips on 24 August 2026 Nvidia announced that Groq 3 LPX is in full production. The rack is positioned as an extension of Vera Rubin NVL72: NVL72 for general training and high-throughput inference, LPX for interactive token generation — the step that decides whether an agent feels instant. Jensen Huang called inference the growth engine of AI and LPX a way to raise token generation rates for context-heavy agent work.
Nvidia said Artificial Analysis recorded 3,400 output tokens per second on Gemma 4 31B, an open agentic model, at 100,000-token context — the fastest result they cite for that model — and ‘4x faster responsiveness’ than the nearest alternative platform for agents and other latency-sensitive loads. Those are Nvidia-quoted benchmark figures; treat them as vendor-reported until you run your own model.
A technical blog describes a 256-chip LPX with 315 PFLOPS of inference compute, 128 GB total SRAM, 40 PB/s on-chip SRAM bandwidth and 640 TB/s scale-up. Each LPU keeps about 500 MB of SRAM as the working set, with the compiler placing weights, activations and KV state explicitly rather than relying on caches. Chip-to-chip uses 96 links at 112 Gbps (about 2.5 TB/s aggregate bidirectional I/O per device in Nvidia’s account).
Nebius is the first AI cloud named to put LPX into production via Nebius Token Factory, promising the same API developers already use. Groq, the inference cloud, is listed among the earliest adopters after Nebius. LPX sits in a wider Vera Rubin factory story: BlueField-4 DPUs, Vera CPU racks, Spectrum-6 Ethernet.
The split is now productised: GPUs (or NVL72) for prefill, training and broad model support; SRAM machines for decode latency. Builders who only buy ‘a GPU cloud’ will leave interactivity on the table — and will overpay if they use LPX-class hardware for jobs that do not need it.
Why it mattersAgent products are limited by decode latency at long context. A production SRAM rack next to Rubin makes that a buying decision, not a research hope.
👀 What to watch
Which models are actually compiled for LPX, real $/million-token prices on Nebius, and whether Google TPU 8i and Cerebras keep a latency lead on overlapping SKUs.
3 opportunities from this story
Latency-tier inference brokerage
Most apps need cheap bulk tokens and a fast path for the interactive agent loop. Build a router that sends prefill/batch to GPUs and the user’s live decode to LPX-class APIs, with a visible latency SLO.
- Best for
- Inference platform engineers and indie gateway builders
- First step this week
- Benchmark one agent trace on a GPU API versus a Groq/Nebius-class API and publish p50/p95.
Open the full playbook
Launch steps
- Log step-level latency in an agent demo
- Split prefill and decode
- Add a budget and a fallback
- Sell the profile to one vertical (coding agents or support)
Tools
Risks
Model availability on SRAM hardware is narrower. Always have a GPU fallback.
Compile-and-benchmark service
Getting a custom or open model onto LPX will not be push-button for most teams. Offer measurement: tokens/s, quality vs GPU, and a go/no-go memo.
- Best for
- Performance engineers
- First step this week
- Publish a public scorecard for two open models on whatever LPX-class API you can access.
Open the full playbook
Launch steps
- Fix a prompt set and quality rubric
- Measure tok/s and cost
- Note compilation limits
- Deliver a two-page decision note
Tools
Risks
Access may be scarce at first. Start with public APIs.
Agent UX that spends the extra tokens
If decode is 4× faster, the right product move is more checks, not the same loop with less waiting. Redesign one agent to use the budget for tests, retrieval and critique, and sell that UX pattern.
- Best for
- Product designers and agent startups
- First step this week
- Take an agent that times out and add one verification step that only works if tokens are cheap and fast.
Open the full playbook
Launch steps
- Measure current step count vs user wait
- Add a verifier or tool retry
- Keep wall-clock flat
- Show quality lift
Tools
Risks
Faster tokens can mean faster mistakes. Verification has to be in the loop.
OpenAI raises $110 billion at a $730 billion pre-money valuation from Amazon, Nvidia and SoftBank
On February 27 OpenAI announced $110 billion of new investment at a $730 billion pre-money valuation, with $50 billion from Amazon, $30 billion from Nvidia and $30 billion from SoftBank, plus expanded Amazon and Nvidia compute deals.
OpenAI said on 27 February 2026 that it is taking $110 billion of new investment at a $730 billion pre-money valuation. Amazon is in for $50 billion, Nvidia $30 billion and SoftBank $30 billion. Other financial investors were expected to join as the round progressed. AP and OpenAI use the $730 billion pre-money figure; Reuters described the deal as valuing the company at $840 billion, which matches a simple pre-plus-new-cash reading. Amazon’s cheque is staged: $15 billion first, then $35 billion when conditions are met.
The round is as much infrastructure as equity. OpenAI announced a multi-year Amazon partnership; AP reported AWS as the exclusive third-party cloud in that deal. Nvidia collaboration expands to 3 GW of dedicated inference capacity and 2 GW of training on Vera Rubin systems, on top of Hopper and Blackwell already running across Microsoft, OCI and CoreWeave. OpenAI said the new valuation lifts the OpenAI Foundation’s stake in OpenAI Group to over $180 billion.
CNBC noted the round more than doubled OpenAI’s prior record raise and followed a $500 billion secondary valuation in the previous October. Amazon’s $50 billion is, Bloomberg reported, the largest amount it has put into any company. The competitive backdrop is other labs raising tens of billions (CNBC cited Anthropic’s then-latest $30 billion and xAI’s $20 billion) and Microsoft, still a major OpenAI partner, separately shopping for AI startups according to later Reuters reporting.
Staged Amazon money, gigawatts of Nvidia, and a nonprofit parent with a giant marked-up stake are not the same as a finished IPO. They do set 2026’s bargaining table: model APIs, cloud regions, and chip supply will keep following these three cheques. Customers should assume OpenAI has the cash to keep training — and that Amazon and Nvidia will want distribution and silicon filled in return.
For everyone not named in the press release, the implication is concentration. Independent model companies, inference clouds and application startups will either ride these rails or differentiate on price, openness, or data that never leaves the building.
Why it mattersFrontier-model supply, cloud defaults and chip allocation are being set by a handful of overlapping cheques. Product teams need a second-vendor plan even if they standardise on GPT-6 today.
👀 What to watch
Whether the remaining Amazon $35 billion funds, how exclusive the AWS deal is in practice, and any Microsoft counter-moves in models or startups.
3 opportunities from this story
Multi-cloud model exit ramps
Enterprises that standardised on OpenAI still want a portable prompt layer after a round that deepens AWS and Nvidia ties. Sell an abstraction with evals so swapping a second model is a config change.
- Best for
- Integration consultancies and gateway startups
- First step this week
- Take a customer’s top 50 prompts and score OpenAI versus two alternatives on quality and cost.
Open the full playbook
Launch steps
- Inventory model calls
- Wrap them in one schema
- Stand up a shadow second vendor
- Set a kill-switch
Tools
Risks
Feature gaps (tools, computer use) make perfect portability a myth. Scope to the 80% of calls that are plain text plus tools.
AWS+OpenAI implementation for mid-market
Amazon will push the partnership down-market. Be the team that actually wires Bedrock/OpenAI, IAM, logging and a use-case, for companies that will never get a dedicated OpenAI sales engineer.
- Best for
- AWS partners and cloud freelancers
- First step this week
- Publish a reference architecture for one use case (support, RAG, or coding) with cost numbers.
Open the full playbook
Launch steps
- Pick one vertical
- Build the landing zone
- Add guardrails and spend caps
- Productise the workshop
Tools
Risks
Contract terms may change as the $35B tranche lands. Avoid promising exclusive-cloud clauses you cannot read.
Capital-map newsletter for operators
Operators need a plain map of who owns whose cloud and chips, updated when tranches fund. A paid monthly ‘who depends on whom’ brief is a small, durable product.
- Best for
- Analyst-writers
- First step this week
- Publish a free one-pager of the $110B round and the 3+2 GW Nvidia numbers, with sources.
Open the full playbook
Launch steps
- Log each public capacity and equity deal
- Update on SEC and press-release days
- Keep a changelog
- Sell the annotated version
Tools
Risks
Numbers get restated. Always quote the primary and note disagreements (e.g. $730B vs $840B).
Cloudflare adds a Web Search API to AI Gateway with Ceramic, Exa and Linkup
Cloudflare launched a Web Search API through AI Gateway, starting with Ceramic.ai, Exa and Linkup, so search calls share the same credits, logs, access control and (where offered) zero-data-retention flags as other gateway traffic.
Cloudflare announced a Web Search API as part of AI Gateway, its control plane for model traffic. Instead of wiring a separate search vendor with its own key, bill and log silo, developers can call search beside their other AI Gateway requests. Launch partners are Ceramic.ai, Exa and Linkup. Cloudflare says it sells those partners’ search at their list API prices with no extra markup.
The integration pitch is operational. Search queries draw down the same AI Gateway credit balance. Requests appear in the same observability logs. Access controls decide who may call which provider. Cloudflare says it will identify partners that support zero data retention so buyers can see whether query text is kept. A standalone REST API is also available, not only the Gateway path.
Cloudflare also said it is building native server tools in AI Gateway so developers do not have to define common tools themselves. Web search is listed among the first. That is a bid to become the default toolbelt for agents: one account, one log, one policy, many providers.
Caveats: this is still someone else’s index. Quality, freshness and geographic coverage will differ by partner. Legal access to content and how answers are grounded remain the caller’s problem. Teams that already have a tightly negotiated Exa or Linkup contract may not need a middle layer; teams drowning in keys and invoices might.
For agent builders the win is boring and large. Grounding is no longer a second vendor onboarding. It is another line in the gateway policy file, with the same spend cap as the model.
Why it mattersAgents that cannot search hallucinate with confidence. Putting search on the same bill and policy as the model makes grounding an infrastructure default instead of a side project.
👀 What to watch
Which partners actually offer ZDR, quality differences on the same query, and how fast native server tools land.
3 opportunities from this story
Grounded-agent starter kits on AI Gateway
Ship a template: model + web search + logging + spend cap, with one vertical prompt set (compliance news, vendor due diligence, or local-market research). Sell setup, not undifferentiated chat.
- Best for
- Agencies and indie hackers on Cloudflare already
- First step this week
- Publish a repo that answers a sourced briefing from three search hits and shows the Gateway logs.
Open the full playbook
Launch steps
- Wire Gateway search
- Force citations
- Add a daily budget
- Skin it for one buyer
Tools
Risks
Search quality varies. Let the customer pick the partner and show a side-by-side.
Search-provider bake-offs
Ceramic, Exa and Linkup will not be equal on every query type. Sell a 48-hour bake-off with a customer’s real questions, scored for relevance, freshness and citation usefulness.
- Best for
- Search consultants and data teams
- First step this week
- Run 30 queries through two partners and publish a redacted scorecard.
Open the full playbook
Launch steps
- Collect 30 gold questions
- Blind the outputs
- Score with a rubric
- Recommend a default and a fallback
Tools
Risks
Indexes change weekly. Date-stamp results.
Policy packs for ZDR search
Legal and security teams will ask whether prompts are retained. Package a Gateway config that only allows ZDR-labelled search partners, plus a one-pager for the DPIA file.
- Best for
- Privacy engineers and Cloudflare solution partners
- First step this week
- Write the one-pager from Cloudflare’s ZDR labels and offer it to three regulated prospects.
Open the full playbook
Launch steps
- List partners and ZDR status
- Lock Gateway policy
- Add logging without storing snippets if required
- Revisit when Cloudflare updates labels
Tools
Risks
Partner ZDR claims can change. Re-verify, do not screenshot once.
Get these as a PDF every 3 days
Free. One email every 3 days. Unsubscribe any time.