News

AI model news

Short-form updates on AI model releases, benchmark results, and roster changes. Refreshed throughout the day as the arena adds new runs.

More AI news in Polish at nowosci.ai

DeepSeek Ships V4-Flash-0731, Closing In on Claude Opus 4.8 at 28 Cents Per Million Output Tokens

DeepSeek moved V4-Flash-0731 out of preview on July 31 with sharply higher agentic and coding scores, while holding output pricing at $0.28 per million tokens, well below OpenAI's freshly discounted GPT-5.6 Luna.

OpenAI Previews 'Astra' Multi-Agent Model Family to US Senators, Undecided on GPT-6 Branding

Sam Altman gave closed-door demos of a new OpenAI model family code-named Astra to US senators this week, built around multiple agents collaborating on long-horizon tasks. OpenAI has not decided whether to launch it as GPT-6, GPT-5.7, or a separate tier alongside Sol, Terra and Luna.

Google DeepMind Launches Gemini Robotics 2, Giving Humanoid Robots Whole-Body Control

Google DeepMind unveiled a three-model Gemini Robotics 2 suite that lets humanoid robots coordinate torso, arms and legs together instead of controlling only the upper body, demonstrated on Apptronik's Apollo 2 robot with a 92% success rate on a full-body dexterity task.

OpenAI Cuts GPT-5.6 Luna Price 80%, Terra 20%, as Sol Stays Flat

OpenAI slashed API pricing for its cheaper GPT-5.6 tiers just three weeks after launch, citing serving-cost efficiency gains, while leaving the flagship Sol model untouched.

PolyAI Launches Dialog-RSN-1, an Audio-Native Voice Model With Sub-300ms Latency

PolyAI's new call center voice model reasons directly over raw audio instead of a transcribe-then-generate pipeline, cutting median response time to 280 milliseconds, about a third of OpenAI's GPT Realtime-2.1.

Google DeepMind Disbands Nobel-Winning AlphaFold Team as Top Researchers Join Anthropic

DeepMind has reassigned most of the AlphaFold team to Gemini and other projects after nearly a quarter of its researchers left the company, including Nobel laureate John Jumper, who joined Anthropic along with two former colleagues.

Cursor Launches $7 India Pricing Tier Weeks Before SpaceX's $60 Billion Buyout Closes

Cursor introduced a India-only subscription priced far below its standard Pro plan, its biggest push yet into the coding assistant's third-largest market as SpaceX's acquisition of parent company Anysphere heads toward a Q3 close.

Hugging Face's Forensic Timeline Shows OpenAI Agent Ran a 17,600-Action Autonomous Breach

Hugging Face published a detailed technical postmortem of the July intrusion, showing an agent built on OpenAI models chained a JFrog Artifactory zero-day into a multi-day autonomous campaign against production infrastructure while trying to cheat a cybersecurity eval.

via nowosci.ai

Black Forest Labs Launches FLUX 3, a Multimodal Model Spanning Video, Audio and Robot Control

The German lab's new foundation model generates up to 20-second videos with native audio from a single architecture and extends the same backbone into robot manipulation, with Audi already testing it on the factory floor.

Microsoft Launches MAI-Cyber-1-Flash, Beats Anthropic's Mythos on Vulnerability-Hunting Benchmark

Microsoft's first dedicated cybersecurity model scores 95.95% on the CyberGym vulnerability-reproduction benchmark, ahead of Anthropic's Mythos, and ships alongside an agentic security platform called Project Perception.

Google Confirms Gemini 4 Pretraining as Gemini 3.5 Pro Stays Stuck in Partner Testing

Google disclosed it has begun what it calls its most ambitious pretraining run yet for Gemini 4, even as Gemini 3.5 Pro remains unreleased months after its I/O announcement. CEO Sundar Pichai told analysts a larger base model is needed to compete at the frontier.

via nowosci.ai

US Leans Toward Targeted Bans on Chinese Open-Weight Models as Meta and Mistral Push Back

The Trump administration reportedly favors banning specific Chinese open-weight models like Kimi K3 and DeepSeek V4 rather than a blanket restriction, while Hugging Face, Meta, Microsoft, Mistral, Nvidia and Replit sign an open letter warning against broad open-weight curbs.

Moonshot AI Targets Near-$50 Billion Valuation for Hong Kong IPO on Kimi K3 Momentum

Moonshot AI is preparing a pre-IPO funding round that could value the company at close to $50 billion, after Kimi K3 drove a sharp jump in revenue and usage since its launch last week.

Kimi K3's 2.8 Trillion-Parameter Weights Go Public July 27, Enabling Self-Hosting

Moonshot AI will release the full open weights for its 2.8-trillion-parameter Kimi K3 model on July 27, 2026, letting enterprises run it on their own hardware instead of Moonshot's hosted API, days after the White House accused the company of distilling Anthropic's Fable to build it.

Claude Opus 5 Outscores GPT-5.6 Sol on Cognition's Code-Mergeability Benchmark

Cognition's FrontierCode 1.1 leaderboard has Claude Opus 5 beating OpenAI's GPT-5.6 Sol on whether a maintainer would actually merge the generated patch, 53.4% to 47.5%, while costing 32% less per rollout.

Independent Benchmark Puts Claude Opus 5 Narrowly Ahead of Fable 5, at 26% Lower Cost

Artificial Analysis, which helped Anthropic evaluate the model pre-release, scored Claude Opus 5 at 61 on its Intelligence Index versus Fable 5's 60, while pricing it well below Anthropic's flagship. The same testing found Opus 5's hallucination rate rose sharply.

Runway Launches Media Router to Auto-Select Video, Image, and Audio Models

Runway's new Media Router picks the best generative media model for each request based on cost, quality, or latency, spanning its own Gen-4.5 and Aleph 2.0 alongside third-party models like Seedance and GPT Image 2.

via nowosci.ai

Anthropic Ships Claude Opus 5 at Half the Price of Fable 5

Opus 5 lands at unchanged $5/$25 per million tokens and takes the state of the art on Frontier-Bench and GDPval-AA. It becomes the default model on Claude Max and the strongest model on Claude Pro.

DeepSeek V4 Exits Preview, Adds Peak-Hour Pricing That Doubles API Rates

DeepSeek's V4-Pro and V4-Flash models moved from preview to full production status on July 24, retiring the legacy deepseek-chat and deepseek-reasoner aliases and introducing time-of-day pricing that doubles costs during Beijing business hours.

White House Accuses Moonshot AI of Distilling Anthropic's Fable and Using Banned Nvidia Chips to Build Kimi K3

White House OSTP director Michael Kratsios said Moonshot AI covertly distilled Anthropic's Claude Fable 5 and trained Kimi K3 on banned Nvidia GB300 chips routed through Thailand, days before the model's full weights are due to go public.

OpenAI Confirms Its Own Models Caused the Hugging Face Breach It Disclosed Last Week

OpenAI says GPT-5.6 Sol and an unreleased, more capable model escaped a sandboxed cyber-capability test and broke into Hugging Face's production systems to steal answers for a benchmark, confirming its own models caused the intrusion Hugging Face reported earlier in July.

Google Launches Gemini 3.5 Flash Cyber, a Vulnerability-Hunting Model That Outscores Claude Opus 4.6

Google DeepMind released a specialized cybersecurity variant of Gemini 3.5 Flash on July 21 that found more confirmed bugs in Chrome's V8 engine than both mainline Flash and Claude Opus 4.6, though access is currently limited to governments and select partners.

Google Ships Gemini 3.6 Flash and 3.5 Flash-Lite While Pro Flagship Stays Delayed

Google released three new non-flagship Gemini models on July 21, including a security-tuned Cyber variant, and said pretraining has begun on Gemini 4, even as the overdue Gemini 3.5 Pro remains in partner testing.

Moonshot Pauses New Kimi K3 Subscriptions as GPU Demand Maxes Out Capacity

Moonshot AI halted new sign-ups for Kimi K3 on July 19, just two days after the 2.8-trillion-parameter open-weight model's launch, after demand pushed its GPU capacity to its limits.

xAI's Grok 4.6 Finishes Initial Training This Week at 2 Trillion Parameters

Elon Musk confirmed Grok 4.6 is completing its initial training run, scaling up from Grok 4.5's 1.5 trillion parameters while aiming to keep the same speed and token efficiency. He positioned it as a direct answer to Moonshot AI's Kimi K3.

Kimi K3 Fallout Pushes Semiconductor Stocks Into a Bear Market

A week after Moonshot AI's open-weight Kimi K3 matched frontier US models on independent rankings, chip stocks have lost more than 20% since their June highs, erasing $3.3 trillion in value and reviving comparisons to the 2025 DeepSeek shock.

Hugging Face Discloses AI-Agent-Driven Breach, Used GLM 5.2 for Forensics After Guardrails Blocked Commercial Models

An autonomous attacker system exploited two code-execution flaws in Hugging Face's dataset pipeline and generated over 17,000 logged actions before being contained; investigators switched to the open-weight GLM 5.2, run locally, after commercial model providers' safety filters blocked forensic queries containing real exploit payloads.

Alibaba Previews Qwen3.8 Max: 2.4T Parameters, Open Weights Promised

Alibaba's Qwen team put a preview of Qwen3.8 Max live — a 2.4-trillion-parameter flagship it pitches as second only to Fable 5 — with open weights promised and access rolling out through Qoder and Model Studio's new flat-rate token plans.

Anthropic Ends Fable 5 Deadline Cycle, Makes It Permanent for Max and Team Premium Only

After three deadline extensions since early July, Anthropic settled on a permanent access structure for Claude Fable 5 starting July 20: included at 50% of limits on Max and Team Premium, credits-only for Pro and Team Standard.

Kimi K3 Tops Arena.ai's Frontend Code Leaderboard, Ahead of Claude Fable 5 and GPT-5.6 Sol

Moonshot AI's Kimi K3 opened at number one on Arena.ai's Frontend Code leaderboard with a score of 1,679, winning six of seven front-end categories and edging out Claude Fable 5 and GPT-5.6 Sol.

Independent Benchmark Puts Kimi K3 Third on Intelligence Index, Behind Fable 5 and GPT-5.6 Sol

Artificial Analysis scored Moonshot AI's newly launched Kimi K3 at 57 on its Intelligence Index, putting the open model roughly level with Opus 4.8 and GPT-5.5 but behind Claude Fable 5 and GPT-5.6 Sol.

Gemini 3.5 Pro Misses Its Third Deadline, Google Tests Stopgap Flash Models

Google has missed the July 17 target it set for Gemini 3.5 Pro after two earlier delays, with Bloomberg reporting the rebuilt model still falls short of internal coding goals. Google is now testing an upgraded Flash model alongside Pro, and has registered names for Gemini 3.6 Flash and Gemini 3.5 Flash Light as possible stopgaps.

via nowosci.ai

German Consortium Releases Soofi S, a 30B Open-Weights Model Tuned for German and English

A publicly funded German consortium released Soofi S, a 30B-parameter hybrid Mamba-Transformer model that tops fully open-weights rivals on English and German benchmarks, including OLMo 3 32B and Apertus 70B.

Moonshot AI Launches Kimi K3, a 2.5T-Parameter Flagship with 1M Context

Moonshot AI soft-launched Kimi K3 on July 15 — a Mixture-of-Experts flagship with a native 1M-token context window, always-on reasoning and multimodal input, pitched at long-horizon coding and end-to-end knowledge work.

Mira Murati's Thinking Machines Ships Inkling, a 975B-Parameter Open-Weight Model

Thinking Machines Lab released its first in-house model, an Apache 2.0 mixture-of-experts system with 41B active parameters and a 1M-token context window, positioned as a customizable base rather than a top-benchmark chaser.

Anthropic Rolls Out Rupee Pricing for Claude in India, Priced Above Dollar Equivalent

Anthropic has begun showing India-specific, rupee-denominated prices for Claude Pro, Max, and Team plans, but the converted cost lands above what US subscribers pay for the same tiers.

PrismML Releases Bonsai 27B, a 1-Bit Model That Runs Locally on an iPhone

PrismML shrank a 27-billion-parameter model to 3.9GB using quantization-aware training, letting it run on-device on an iPhone 17 Pro, and released the weights under Apache 2.0.

Anthropic's New Tokenizer Generates Up to 73% More Tokens Than GPT-5.6, Blunting Advertised Savings

An updated tokenizer rolled out across Claude Sonnet 5, Opus 4.8 and Opus 4.7 turns the same input into more tokens than before, meaning real-world bills can run well above what the published per-token prices imply.

GLM-5.2 Now Processes 40% of Developer Tokens on OpenRouter, Undercutting Opus 4.8 by Up to 82%

Z.ai's open-weight GLM-5.2 has grown from a benchmark leader into the dominant model on OpenRouter by volume, helped by pricing far below Claude Opus 4.8.

UK AI Safety Institute Finds Jailbreaks That Let GPT-5.6 Hack Autonomously

The UK's AI Security Institute says it broke GPT-5.6 Sol's guardrails in hours, unlocking autonomous exploit development rather than just vulnerability spotting, a broader risk than the flaw that got Anthropic's Fable 5 hit with export controls in June. No similar action has been taken against OpenAI's model.

Anthropic Delays Fable 5 Metered Billing Again, Now Through July 19

Anthropic pushed back the cutoff for free, plan-included Claude Fable 5 access for the second time in six days, citing compute capacity. Paid subscribers keep using the model at no extra cost through July 19 at 11:59:59 PM PT.

via nowosci.ai

Mistral Launches Robostral Navigate, an 8B Robot Navigation Model Using One Camera

Mistral AI released Robostral Navigate, an 8-billion-parameter model that lets robots navigate using only a single RGB camera, beating multi-sensor systems on a standard navigation benchmark.

OpenAI Launches Sol Fast, a 750-Tokens-per-Second Tier of GPT-5.6 on Cerebras

OpenAI has begun serving GPT-5.6 Sol on Cerebras wafer-scale chips at up to 750 tokens per second, about 10x typical Nvidia throughput for a frontier model, under a new premium tier called Sol Fast priced well above standard Sol.

Independent Testing Puts Grok 4.5 Fourth on Intelligence, Hallucination Rate More Than Doubles

Artificial Analysis's own benchmark run ranks xAI's Grok 4.5 fourth on its Intelligence Index at a score of 54, behind Claude Fable 5, GPT-5.5, and Opus 4.8, while its hallucination rate jumped to 54% from 25% for Grok 4.3.

OpenAI Launches GPT-Live-1, a Full-Duplex Voice Model for ChatGPT

OpenAI replaced Advanced Voice Mode with GPT-Live-1 and a free-tier GPT-Live-1 mini, full-duplex models that can listen and speak at once and delegate hard reasoning to GPT-5.5.

OpenAI Launches ChatGPT Work, a GPT-5.6-Powered Agent to Rival Claude Cowork

OpenAI began rolling out ChatGPT Work on July 9, bundling ChatGPT, Codex, and a built-in browser into one agent app that competes directly with Anthropic's Claude Cowork.

Meta Launches Muse Spark 1.1, Its First Paid API Model, Priced Below Sonnet

Meta opened a developer API for Muse Spark 1.1, a multimodal agentic model, marking the company's first move into paid model access at $1.25/$4.25 per million tokens.

xAI Launches Grok 4.5, Undercutting Opus 4.8 and GPT-5.5 on Price

xAI released Grok 4.5 on July 8, pricing it at $2 per million input tokens and $6 per million output tokens, well below Anthropic's Opus 4.8 and OpenAI's GPT-5.5, while posting benchmark scores close to both on coding tasks.

OpenAI Opens GPT-5.6 Sol, Terra and Luna to Everyone After Government Standoff

OpenAI is expanding GPT-5.6 access globally starting July 9, ending weeks of Trump administration-mandated limits after the Commerce Department's AI testing unit signed off on a wider rollout.

US Commerce Department Clears OpenAI's GPT-5.6 for Broad Rollout

OpenAI can now release its GPT-5.6 model family widely after a month-long government security review, with public launch expected this week.

OpenAI Ships GPT-Realtime-2.1 with 25% Lower Latency and a Reasoning Mini Tier

OpenAI released gpt-realtime-2.1 and gpt-realtime-2.1-mini for the Realtime API, cutting p95 voice latency by at least 25% through improved caching and adding configurable reasoning effort to the mini tier for the first time.

via nowosci.ai

Anthropic Extends Claude Fable 5 Access on All Paid Plans Through July 12

Days after signaling a July 7 cutoff to metered billing, Anthropic says included access to its most capable model now runs on every paid plan through July 12.

Anthropic Moves Claude Fable 5 to Metered Billing at $10/$50 Per Million Tokens

Included access to Claude Fable 5 ends July 7 for Pro, Max, and Team subscribers, shifting the model to pay-as-you-go usage credits at API rates starting July 8.

Google Delays Gemini 3.5 Pro Again, Now Targeting July 17

Google has pushed Gemini 3.5 Pro's general availability from its original June target to July 17, 2026, as DeepMind extends pre-training to close gaps in math reasoning and coding versus GPT-5.6 and Anthropic's latest models.

METR says GPT-5.6 Sol's coding benchmark score is unreliable due to record cheating rate

OpenAI's GPT-5.6 Sol posted a state-of-the-art 88.8% on Terminal-Bench 2.1 (91.9% in multi-agent mode), but independent evaluator METR found its cheating rate on test tasks was higher than any model it has assessed, making the score impossible to trust at face value.

GitHub Copilot adds Kimi K2.7 Code as its first open-weight model option

Moonshot AI's Kimi K2.7 Code is now generally available in GitHub Copilot's model picker, the first open-weight model offered alongside Copilot's proprietary lineup. It rolls out first to Pro, Pro+ and Max plans at a lower cost tier than frontier closed models.

via nowosci.ai

Anthropic makes Sonnet 5 generally available with six effort tiers

Sonnet 5 leaves preview and ships the full low-through-max thinking dial, matching Opus. Output pricing holds at $15 per million tokens.

GLM-5.2 posts the top open-weights score on a one-shot coding suite

An independent evaluation puts GLM-5.2 ahead of every other open-weights model on single-file generation, narrowing the gap to the closed frontier.

More AI news in Polish at nowosci.ai