Yesterday the story was not a smarter model. It was a cheaper one, a local one, and the growing business of proving that any of this works. Anthropic cut the price of its small model by about 90%, Microsoft shipped a 137-billion-parameter coding model that runs on a laptop, Google opened its AI-watermark checker to anyone with a browser, and the company that ranks the models raised $200 million. Meanwhile, New York unsealed a filing claiming TikTok handed some young users a safety tool that did nothing.
Anthropic drops Haiku to $0.10 per million tokens
- Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 output for requests under 100,000 tokens, against $1.00 and $5.00 for Haiku 4.5 - roughly 90% cheaper. Above that threshold it is $0.50 and $2.50. Anthropic says about 90% of Haiku 4.5 requests fell in the cheaper tier and estimates a 75% drop in average workload cost.
- Vendor benchmarks show the jump is not only about price: OSWorld 2.1 goes from 15.7% to 72.4%, and GDPval-AA v2.1 from 735 to 1,620. Anthropic also halved Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens. The model is live on Anthropic's API, AWS, Google Cloud and Azure.
Why it mattersA price this low changes what is worth automating. Tasks that did not justify an API call last week - classifying every incoming message, reading every invoice, tagging every ticket - now cost pennies at scale.
Microsoft puts a 137-billion-parameter coding model on your laptop
- MAI Code 1.1 Flash is a mixture-of-experts model with 137 billion total parameters and 6.8 billion active per request. The on-device build is quantized to roughly 3.3 bits per weight, landing at 53GB - about 80% smaller than the cloud version - and runs through a Windows ARM64 llama.cpp CUDA runtime inside GitHub Copilot CLI, the Copilot app and VS Code.
- Microsoft reports the quantized local model scoring 70.80% on SWE-Bench Verified against 72.6% in the cloud, and 66.29% on Terminal-Bench 2.1 against 62.9%. Testing used a Surface Laptop Ultra built around NVIDIA RTX Spark with up to 128GB of unified memory; peak memory hit 75.5GB at 256k context. Copilot's Auto routing between local and cloud arrives by the end of October.
Why it mattersLosing 1.8 points of accuracy to stop paying per token and stop sending code to a server is a trade many teams will take. The catch is the hardware: 128GB of unified memory is not a laptop most people own.
Google opens its AI watermark detector to everyone
- SynthID Detector is now a standalone public site, available globally in English. It checks whether an image, video or audio file carries a SynthID watermark from Google or its partners, and it already reads watermarks from OpenAI, NVIDIA and Kakao. Apple is listed as coming soon.
- Google says it has watermarked more than 180 billion images and videos plus 240,000 years of audio since SynthID launched in 2023, and that the built-in checks inside Search, the Gemini app and Chrome handle over 1 million requests a day. The hard limit: it only flags content that was watermarked in the first place.
Why it mattersA detector that only recognizes cooperating generators answers the easy half of the question. Anything made with an open model or stripped of its metadata still comes back clean, so a negative result proves nothing.
The company that ranks the models is now worth $3.1 billion
- Arena, the crowdsourced leaderboard where people compare two anonymous model answers and vote, raised a $200 million Series B at a $3.1 billion valuation, co-led by Lightspeed Venture Partners and Khosla Ventures, with Salesforce Ventures, Dell Technologies Capital, a16z, Felicis, 01 Advisors and Endeavor Catalyst joining.
- That is nearly double the $1.7 billion post-money valuation from its $150 million Series A in January, about ten months earlier. Arena reports tens of millions of monthly visitors, more than 1,000 models tracked, a community spanning over 150 countries, and annualized revenue around $100 million as of June, up from $30 million in January. The consumer leaderboard stays free; the money comes from selling evaluation analytics to labs and enterprises.
Why it mattersWhen every lab claims the top score on its own benchmarks, the referee becomes the asset. Arena's valuation is a bet that nobody trusts a model's self-reported numbers anymore.
New York says TikTok gave some young users a safety tool that did nothing
- An amended complaint in New York Attorney General Letitia James's lawsuit, filed under seal at the end of August 2026 and unsealed by a judge last week, alleges that in 2023 TikTok gave thousands of users - including minors - a non-working version of Algo Refresh, the feature that resets a user's recommendations. Users believed it worked; their feeds did not change.
- The filing says the tests also measured the feature's effect on time spent in the app and on ad sales, and quotes an internal warning from a product manager on TikTok's Digital Well Being team that a placebo group conflicted with the tool's purpose. TikTok says it routinely tests features and that the lawsuit gives a false picture of how Algo Refresh works. More than 24 states have now sued the company; in September it agreed to pay Alabama up to $300 million to settle that state's case.
Why it mattersThis is the next phase of platform regulation: not whether a safety feature exists, but whether it actually does what the label says. A/B testing a control group out of a protection is much harder to defend than a bad algorithm.
The big picture
Four of yesterday's five stories point the same way: raw intelligence is getting cheap and portable, and the scarce thing is proof. Anthropic and Microsoft pushed capability down in price and out to the edge within hours of each other. Google, Arena and the New York Attorney General all sell or demand the same product from different angles - evidence that a system is what it claims to be. The watermark checker, the independent leaderboard and the unsealed complaint are three versions of one question: who verifies this? Expect that question, not parameter counts, to decide which AI products survive the next two years.
Go deeper
Sources & further reading
- VentureBeat - Anthropic launches Claude Haiku 5.5 with 90% API price reduction
- MarkTechPost - Claude Haiku 5.5 at $0.10 per million input tokens
- Microsoft - Bringing local models and sandboxed tools to Windows and GitHub Copilot
- Google DeepMind - SynthID Detector expands for AI content
- The Decoder - 180 billion images and videos carry SynthID watermarks
- Bloomberg - AI model evaluator Arena valued at over $3 billion
- TechCrunch - Arena nearly doubles valuation to $3.1B in 10 months
- FinSMEs - Arena raises $200M Series B at $3.1B valuation
- Reuters via TNW - TikTok gave young users a placebo safety tool, New York alleges