On Friday, Anthropic launched Opus 5, the latest iteration of its heavyweight model. Arriving just two months after Opus 4.8 and following the June releases of Mythos 5, Fable 5, and Sonnet 5, Opus 5 rounds out the new generation — only the lightweight Haiku waits for an upgrade. While smaller than the flagship Fable 5, Opus 5 is significantly cheaper and noticeably less restrictive. Opus 5 is designed to work on programming tasks for much longer without constant human input.Opus 5 is designed to work on programming tasks for much longer without constant human input. Benchmarks at bargain prices Priced at $5 per million input tokens and $25 per million output tokens, Opus 5 upends the cost-to-performance ratio for agentic tasks. On coding and knowledge work evaluations like Frontier-Bench and GDPval-AA, Opus 5 establishes a new state of the art. In the OSWorld 2.0 computer use benchmark, it outperforms every other model at any given cost, surpassing Fable 5’s best result at just over a third of the price. On ARC-AGI 3, an evaluation where the model has to solve novel problems, its score is three times as high as the next best model.Because Opus 5 is less expensive to run, teams can afford to let it work through larger coding tasks. That also means thinking differently about security.Anthropic writes in the company announcement that Opus 5 is “much stronger at verifying its work and iterating carefully until it succeeds.” During benchmark testing, the model was given an incomplete prompt and intentionally prevented from viewing a drawing of a machine part. Instead of giving up, it wrote its own computer vision pipeline to reconstruct the part from the image data.Because Opus 5 is less expensive to run, teams can afford to let it work through larger coding tasks. That also means thinking differently about security, which is one reason microVMs are getting more attention. Unlike long-running containers, they give platform teams an isolated environment they can tear down as soon as a task is finished.Security for unsupervised agents When a model is working across multiple systems without constant human oversight, access has to be temporary and easy to revoke. That’s why more attention is shifting to short-lived, just-in-time credentials, similar to the delegated authentication model 1Password is developing for Claude.There’s a catch to giving an AI relentless persistence: it doesn’t know when to quit. Your standard API logs won’t flag that behavior in time. Teams will have to build smarter telemetry that actually understands the context of the workflow. You need a semantic circuit breaker to pull the plug before the model burns through your compute budget.Controlling runaway token spend Opus 5 drops the per-token price tag compared to Fable 5, but autonomous coding can still get expensive fast. An agent working through background loops for hours will churn through millions of tokens before handing off a PR. For platform teams, managing spend by the session becomes the new operational requirement.To keep those background jobs from crashing on edge cases, Anthropic is launching Automatic Fallbacks in beta. If a prompt trips a safety classifier on Opus 5, the API silently reroutes the task to Opus 4.8 instead of throwing a hard error and killing the entire pipeline.Opus 5 also inherits its predecessor’s zero-retention posture, bypassing the 30-day data logging mandatory for Fable and Mythos.To understand where Anthropic expects these environment harnesses to operate, it helps to look at how they are artificially restricting the model. The company is threading a careful needle between short-horizon capability and long-horizon safety. If a prompt trips a safety classifier on Opus 5, the API silently reroutes the task to Opus 4.8 instead of throwing a hard error and killing the entire pipeline.Biological capability, careful limitsAnthropic spokesperson tells The New Stack, Opus 5 is a “meaningful improvement over Opus 4.8 on biology tasks, making it our most capable generally available model for scientific research.” It scores 10.2 percentage points higher than Opus 4.8 on internal organic chemistry benchmarks (like inferring molecular structures from spectroscopy data) and 7.7 percentage points higher in predicting how protein sequence variations determine function.However, Anthropic’s testing revealed “significant limitations on long-running autonomous research tasks,” which they view as carrying the highest potential risk. As a result, Opus 5 retains a similar portfolio of biological safeguards to Opus 4.8. While biology-related requests blocked on Fable 5 will now route to Opus 5 instead of 4.8, the uncapped Mythos 5 remains the stronger model for long-horizon, open-ended work like autonomous drug design campaigns.Opus 5 can review source code for defensive security work. Anthropic says those safeguards should activate far less often than they do with Fable 5 — about 85% less, according to the company.Clearly, the models are improving quickly, but the bigger challenge now is making sure the platforms they’re running on can keep up.Major AI model releases: a timelineRelease dateAI modelCompanyThe New Stack reportJuly 24, 2026Opus 5AnthropicAnthropic’s Opus 5 is almost Fable 5July 21, 2026Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash CyberGoogleGoogle ships 3 new Gemini models. Just not the one everyone’s waiting for.July 16, 2026Kimi K3Moonshot AIKimi K3 tops Arena’s coding leaderboard — and it’s open-weightJuly 9, 2026GPT-5.6 Sol, Terra and LunaOpenAIOpenAI’s GPT-5.6 is now liveJuly 9, 2026Muse Spark 1.1MetaMeta debuts Muse Spark 1.1 — its first paid AI modelJuly 8, 2026Grok 4.5SpaceXAI“Opus-class, but faster”: What Elon Musk says about beating AnthropicJune 30, 2026Claude Sonnet 5AnthropicAnthropic Sonnet 5: It closes the gap with Opus 4.8, and is cheap until AugustJune 9, 2026Claude Fable 5 and Claude Mythos 5 — Mythos restrictedAnthropicAnthropic launches Claude Mythos/Fable 5, but you better try it soonMay 28, 2026Claude Opus 4.8AnthropicClaude Opus 4.8 is here: effort controls, dynamic workflows, cheaper fast mode, better honesty, less deceptionMay 19, 2026Gemini 3.5 FlashGoogleGoogle’s Gemini 3.5 Flash beats the frontier modelsApril 23, 2026GPT-5.5 and GPT-5.5 ProOpenAIOpenAI launches GPT-5.5, calling it “a new class of intelligence”April 16, 2026Claude Opus 4.7AnthropicClaude Opus 4.7 arrives with better vision, memory, and instruction-followingApril 7, 2026Claude Mythos Preview — restricted releaseAnthropicAnthropic’s Claude Mythos is now available, but not for youMarch 5, 2026GPT-5.4 Thinking and GPT-5.4 ProOpenAIOpenAI launches GPT-5.4 Thinking and ProFeb. 19, 2026Gemini 3.1 ProGoogleGoogle’s Gemini 3.1 Pro is mostly greatFeb. 17, 2026Claude Sonnet 4.6AnthropicAnthropic’s new Claude Sonnet 4.6 promises Opus-level coding at Sonnet pricingFeb. 5, 2026GPT-5.3-CodexOpenAIOpenAI’s GPT-5.3-Codex helped build itselfFeb. 5, 2026Claude Opus 4.6AnthropicAnthropic debuts Opus 4.6 with standout scores for solving hard problems that other AIs missDec. 17, 2025Gemini 3 FlashGoogleGoogle’s New Gemini 3 Flash Rivals Frontier Models at a Fraction of the CostNov. 24, 2025Claude Opus 4.5AnthropicAnthropic’s New Claude Opus 4.5 Reclaims the Coding CrownNov. 19, 2025GPT-5.1-Codex-MaxOpenAIOpenAI Says Its New Codex-Max Model Is Better, Faster and CheaperNov. 18, 2025Gemini 3 ProGoogleGoogle Launches Gemini 3 ProSept. 29, 2025Claude Sonnet 4.5AnthropicAnthropic Launches Claude Sonnet 4.5Sept. 15, 2025GPT-5-CodexOpenAIOpenAI Launches a New GPT-5 Model for Its Codex Coding AgentAug. 7, 2025GPT-5OpenAIGPT-5: A Choose Your Own Adventure for Frontend DevelopersMay 22, 2025Claude Opus 4 and Claude Sonnet 4AnthropicAnthropic Launches Its Most Powerful Models for Coding YetApril 16, 2025o3 and o4-miniOpenAIOpenAI Releases New Models Trained for DevelopersApril 14, 2025GPT-4.1, GPT-4.1 mini and GPT-4.1 nanoOpenAIOpenAI Releases New Models Trained for DevelopersThe post Opus 5 costs a third of the price — and that’s actually the problem appeared first on The New Stack.