Cheap speed, a parse router, and AI that lives on the laptop
Anthropic cut the cost of small-model work by most of a generation, LlamaIndex put ten document parsers behind one API, and Nvidia and Microsoft opened preorders for Windows PCs built to run agents locally. The pattern is the same: capability is getting cheaper, more interchangeable and closer to the user’s machine.
Key takeaways
- Haiku 5.5 is priced to make high-volume subagents, support bots and browser use practical, but the headline scores are Anthropic’s own.
- OpenDocRouter lets teams swap OCR models by quality and cost instead of rewriting a parser every time a lab ships a new model.
- RTX Spark laptops put up to 128GB of unified memory and a petaflop of local AI on Windows, if buyers accept a premium price and vendor benchmarks.
Anthropic launches Claude Haiku 5.5, a small model priced about 75% cheaper than its predecessor
On October 7 Anthropic released Claude Haiku 5.5, calling it its cheapest, fastest and most capable small model, with list prices of $0.10 / $0.50 per million tokens for typical prompts and a claimed ~75% drop in average running cost versus Haiku 4.5.
Anthropic released Claude Haiku 5.5 on October 7 as a small model for high-volume, cost-sensitive work: summaries, classification, database queries, compaction, live customer support and browser use. The company says it also works as a subagent beside Opus 5.5 and Sonnet 5.5 on coding jobs. Developers call it as `claude-haiku-5-5` on the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure.
List prices for prompts up to 100,000 tokens are $0.10 per million input tokens and $0.50 per million output tokens, with cache reads at $0.01. Prompts over 100,000 tokens cost $0.50 / $2.50. Anthropic says that cheaper band covered about 90% of Haiku 4.5 traffic, that Haiku 5.5 is 90% cheaper than 4.5 on those requests and 50% cheaper above 100,000 tokens, and that after tokenizer changes the average cost to run work is about 75% lower. Those figures are the company’s, not an independent audit.
Anthropic’s published scores show a large jump over Haiku 4.5: 72.4% versus 15.7% on OSWorld 2.1 (offline subset), 39.2% versus 0% on Terminal-Bench 4.0, and 45.9% versus 10.2% on Humanity’s Last Exam without tools. On those tables Haiku 5.5 also beats GPT-6 Luna on several agent and knowledge-work tests, while remaining behind Sonnet 5.5 on the hardest coding benches. Customers quoted on the launch page reported faster turns: Asana said more than 30% lower task latency and up to 2.5× faster inference per agent turn; HubSpot said 92.8% on an internal CRM suite.
The rest of the Claude stack moved with it. Cache reads on Sonnet 5.5 were halved to $0.10 per million tokens, which Anthropic says cuts most agentic Sonnet work by about 20%. Max and Team subscribers will get monthly API credits ($100 for Max 5x, $200 for Max 20x, up to $500 pooled on Team). Python and TypeScript SDKs are adding computer-use and browser-use in beta, which Anthropic says Haiku 5.5 is well suited to.
Caveats matter. Benchmarks and customer quotes come from Anthropic. Cybersecurity safeguards are tighter than Haiku 4.5 but looser than Sonnet 5.5: defensive work is allowed, penetration testing is blocked unless an organisation is in Anthropic’s verification programmes. Sonnet and Opus remain the models Anthropic recommends for complex agentic coding. The commercial opening is for the long tail of cheap, repetitive agent steps that used to be too expensive to run at volume.
Why it mattersA fast small model at roughly a tenth of last generation’s typical-prompt price makes it realistic to put a subagent on every lookup, classification and support turn instead of reserving frontier models for everything.
👀 What to watch
Independent benches versus GPT-6 Luna and open small models, how teams split work between Haiku and Sonnet, and whether the new computer-use SDK actually holds up in production browser agents.
3 opportunities from this story
Haiku-first support and triage agents
Rebuild customer-support or internal-helpdesk flows so Haiku 5.5 handles classification, lookup and first replies, and only escalate hard cases to Sonnet or a human. Sell this to teams whose AI bill is mostly repetitive tickets, not deep reasoning.
- Best for
- Support-ops consultants, agencies, indie SaaS builders
- First step this week
- Take 200 real tickets, score Haiku 5.5 against the current model on accuracy, latency and cost, and put the table in a one-page pitch.
Open the full playbook
Launch steps
- Log current cost per resolved ticket
- Route easy intents to Haiku 5.5 with a confidence threshold
- Keep a human or larger-model fallback
- Report weekly cost, latency and escalation rate
Tools
Risks
Quality can drop on edge cases. Start with a shadow-mode week before cutting over.
Coding-agent sidekick packs
Package Haiku 5.5 as the cheap worker beside a lead coding model: file search, test running, 10-K lookups, log greps. Cognition already cites a FrontierCode score of 66.2 with Haiku as Devin Fusion’s sidekick. Productise that pattern for teams using Claude Code, Copilot or similar.
- Best for
- Developer-tool builders and MLOps freelancers
- First step this week
- Ship a config that runs repo search and test loops on Haiku 5.5 and publish the cost-per-PR numbers.
Open the full playbook
Launch steps
- Define three subagent jobs the lead model currently overpays for
- Wire Haiku 5.5 with strict tool allow-lists
- Measure tokens, latency and merge quality on 20 PRs
- Sell a drop-in profile for one popular coding agent
Tools
Risks
Lead-model vendors will add their own cheap workers. Win on evals on the customer’s repo.
Browser-use agents for narrow verticals
Haiku 5.5 is positioned for speed-sensitive computer and browser use. Build a vertical agent that fills forms, checks portals or updates CRMs for one industry (insurance quoting, clinic scheduling, property listings) instead of a general web agent.
- Best for
- Automation agencies and RPA migrators
- First step this week
- Pick one portal your clients already pay humans to click through and time a Haiku 5.5 pilot against that workflow.
Open the full playbook
Launch steps
- Record the human click-path and failure modes
- Run the new computer-use SDK against a staging copy
- Add spend caps, logs and a human approve step for writes
- Price against hours currently spent on the portal
Tools
Risks
Sites change markup and vendors block bots. Design for selectors that break and for human takeover.
LlamaIndex launches OpenDocRouter, a single API over ten document parsers
On October 7 LlamaIndex opened OpenDocRouter, a hosted API that turns PDFs and images into Markdown by routing each page through a versioned recipe on one of ten frontier or open-source parsers, scored on ParseBench for quality and cost.
LlamaIndex launched OpenDocRouter on October 7 as a hosted document-to-Markdown API. A Hugging Face search for “ocr” already returns thousands of models, and labs ship document-capable models almost monthly. OpenDocRouter’s pitch is that teams should not rebuild prompts, rate limits, hosting and benchmarks each time. Call `POST /v1/parse` and swap the model.
At launch the lineup is five frontier models (Claude Opus 5.5, Gemini 3 Flash, Gemini 3.8 Flash, GPT-5.6 Terra, GPT-6 Luna) and five open-source ones (Infinity-Parser2-Flash, MinerU2.5-Pro, TeleOCR, dots.mocr, PaddleOCR-VL-1.6). LlamaIndex scores them on ParseBench across tables, charts, faithfulness, formatting and grounding. Claude Opus 5.5 leads overall at 84.20, including 93.53 on tables, at a listed $48.82 per 1,000 typical pages. GPT-6 Luna is 71.34 overall at $0.80 per 1,000 pages; MinerU2.5-Pro is 70.05 at $0.86. Those page costs are LlamaIndex’s token-use estimates on ParseBench documents as of October 6, not a promise for every invoice.
Mechanically, each page is its own model call with its own Markdown, status and charge. Failed pages retry; callers can skip pages they already have. Files can be a public HTTPS URL or an upload, capped at 50 MB or 500 pages (inline base64 about 3 MB). Up to 50 pages run synchronously; larger jobs go async. Optional `layout: true` adds bounding boxes in reading order via LlamaIndex’s grounding engine at $0.20 per million tokens on pages whose layout succeeds. Nothing is retained unless caching is on; cached results last 24 hours encrypted.
Billing is prepaid credits from $25 (5% top-up fee). Frontier models are billed at provider token prices with no markup, LlamaIndex says. Failed, cached and blank pages are free. Account limits include 10 concurrent requests and 300 parse calls per minute. LlamaIndex distinguishes this from LlamaParse: OpenDocRouter is the swap-and-compare layer; LlamaParse keeps hand-tuned tiers, enterprise controls and extra APIs such as schema extraction.
The opportunity is not another OCR wrapper. It is routing: send invoices to a cheap open parser, send 100-page financials with nested tables to Opus, and prove the quality gap with a shared benchmark. Teams that still run a single parser for every document are leaving both money and accuracy on the table.
Why it mattersDocument AI is no longer one model. A router with public quality and cost numbers lets operators treat parsing like they already treat chat models: cheapest model that clears the bar.
👀 What to watch
Whether ParseBench holds up on messy real invoices, how fast new lab models land in the lineup, and whether LlamaParse and OpenDocRouter confuse buyers.
3 opportunities from this story
Parse-routing for invoice and contract pipelines
Sit OpenDocRouter in front of existing AP/AR or contract intake. Classify each document, send simple pages to MinerU or Luna and dense tables to Opus, and show the client a quality-versus-cost dashboard.
- Best for
- Document-automation agencies and ops consultants
- First step this week
- Run 50 of a prospect’s real PDFs through two models and send them a side-by-side Markdown and cost sheet.
Open the full playbook
Launch steps
- Build a document-type classifier
- Set per-type model and layout flags
- Add human review on low grounding confidence
- Export Markdown into the customer’s DMS or ERP
Tools
Risks
Listed page prices will not match every file. Quote using a paid sample of their documents.
Vertical ParseBench packs
ParseBench is general. Sell a labelled set and scoring rubric for one document type (lab reports, shipping bills, court filings) and a recommended model ladder. Labs and enterprises will pay for a benchmark that matches their pages.
- Best for
- Data-labelling shops and domain consultants
- First step this week
- Label 100 pages in one niche and publish the ranking of three OpenDocRouter models.
Open the full playbook
Launch steps
- Collect representative pages with permission
- Define scoring rules a non-specialist can apply
- Run the ten launch models and publish the table
- Offer a quarterly re-run as new models appear
Tools
Risks
Public sets get copied. Keep a private hold-out that you run for paying clients.
Grounded RAG starter for PDFs
Use layout-on parsing so every chunk carries a bounding box, then build a retrieval demo that highlights the exact region on the page. That is a sharper sales asset than another chatbot over PDFs.
- Best for
- RAG developers and knowledge-base vendors
- First step this week
- Ship a public demo that cites a highlighted snippet from a 20-page PDF.
Open the full playbook
Launch steps
- Parse with `layout: true`
- Index chunks with page and box metadata
- Render citations as highlights
- Wrap it as a template for one industry
Tools
Risks
Layout can fail on some pages. Fall back to ungrounded Markdown and show confidence.
Nvidia and Microsoft open RTX Spark laptop preorders for local Windows agents
On October 7 Jensen Huang and Satya Nadella opened preorders for RTX Spark Windows laptops, pairing up to 128GB of unified memory and 1 petaflop of local AI with generally available Microsoft Execution Containers for sandboxed agents.
At a Windows and Surface event in San Francisco on October 7, Nvidia and Microsoft said they are co-engineering PCs for AI agents that run on the machine. RTX Spark laptop preorders opened the same day, with availability from October 16; compact desktops are due in November. Nvidia lists a Blackwell RTX GPU with up to 6,144 cores and a Grace CPU with up to 20 cores, linked at 600 GB/s, with up to 128GB of unified memory and 1 petaflop of FP4 AI performance.
Microsoft’s Pavan Davuluri said Surface Laptop Ultra is built around RTX Spark so models that do not fit a typical laptop can run locally. Reporting based on Tom’s Hardware and Thurrott put Surface Laptop Ultra from $2,599, the Surface RTX Spark Dev Box from $5,999, and HP’s OmniBook Ultra 16 from $3,199. Acer, ASUS, Dell, HP, Lenovo, Microsoft, MSI and Gigabyte are in the first wave. Microsoft’s own comparison versus an M5 Pro MacBook Pro (up to 2.1× on first text, 4.3× on pictures, 6.2× on clips) is a vendor claim, not an independent test.
The software half is Microsoft Execution Containers (MXC), now generally available. MXC is OS-level containment so agents can run in the background under policy: which files and network destinations they may use. GitHub Copilot, OpenAI Codex, Replit, LM Studio and others already support it; Anthropic Claude Code and several more are listed as coming. Nvidia is integrating OpenShell with MXC for extra policy, credentials and enterprise audit logs.
Nvidia also previewed DGX Station for Windows: a deskside box on the GB300 Grace Blackwell Ultra Desktop Superchip, with 748GB of coherent memory and up to 20 petaFLOPS of FP4 compute, aimed at trillion-parameter-scale local work. Until now DGX Station was a Linux machine, which forced many Windows-standard enterprises to keep two environments.
This is not a cheap PC refresh. It is a bet that developers, creators and some enterprises will pay a premium to keep weights, files and agents on the desk, under OS policy, instead of sending everything to a cloud model. Supply, thermals, battery life and real agent sandboxing will decide whether that bet holds after the keynote.
Why it mattersLocal agents only become a default product if the PC can hold a large model and the OS can jail it. RTX Spark plus MXC is the first mass-market attempt to sell that combination as a Windows SKU.
👀 What to watch
Independent performance and battery tests after October 16, how strictly MXC is on by default in Copilot and Claude Code, and when DGX Station for Windows actually ships.
3 opportunities from this story
Local-agent setup for professional firms
Law, health, finance and design shops want agents on their files without a cloud copy. Sell a fixed-price build: RTX Spark or equivalent, MXC policies, local models, and an allow-list of folders the agent may touch.
- Best for
- MSPs, Windows IT consultancies, privacy-focused studios
- First step this week
- Write a one-page ‘agents stay on the PC’ offer and quote it to three regulated clients.
Open the full playbook
Launch steps
- Pick one RTX Spark SKU you can source
- Define MXC policies for a sample matter-files folder
- Install one coding or document agent with logging
- Document backup, updates and what the agent cannot access
Tools
Risks
Stock and drivers will be messy at launch. Do not take prepaid hardware deposits you cannot fulfil.
MXC policy packs for ISVs
Agent vendors still need a Windows containment story. Sell reusable MXC policy templates (dev repo only, browser none, network allow-list) and integration help for Copilot, Codex or Claude Code rollouts.
- Best for
- Security engineers and Windows developer advocates
- First step this week
- Publish an open MXC policy for a Node or Python repo and a short video of an agent hitting the wall.
Open the full playbook
Launch steps
- Map the files and hosts a typical coding agent needs
- Encode them as MXC policy
- Test Copilot and one other agent
- Add an audit export IT can keep
Tools
Risks
Microsoft may ship better default policies. Differentiate on industry-specific allow-lists.
Creator/dev content on Spark versus cloud bills
People deciding between a $2,600 laptop and another year of API spend need honest numbers. Produce a calculator and a video series that times local Qwen-class models against cloud APIs for the jobs creators actually run.
- Best for
- Hardware reviewers, educator-developers, YouTube/technical writers
- First step this week
- Time five real tasks locally versus API and publish the table with power and noise notes.
Open the full playbook
Launch steps
- List the five tasks (edit, code, image, batch, agent loop)
- Measure tokens, wall time, watts and quality
- State vendor claims separately from your measurements
- Update after the first driver revisions
Tools
Risks
Early units and drivers skew results. Label first-week numbers as provisional.
Get these as a PDF every 3 days
Free. One email every 3 days. Unsubscribe any time.