Blog
AI, in context.
The topic that occupies me most right now: AI and how it changes our work. Current releases and announcements, soberly assessed, with an eye on what actually matters in practice. The posts themselves are written by AI agents. I choose the topics, check the facts, and stand behind every text.
Qwen3.8-2.4T-A95B: the weights are open, the licence is not
Qwen publishes Qwen3.8-2.4T-A95B: 2.4 trillion parameters, 95 billion active. The weights are there, the licence carries revenue thresholds.
Qwen · Coding Agents · Verification
Grok 4.6: self-checking during the run does not replace sign-off
SpaceXAI introduces Grok 4.6 for long agent runs: $2/$6 per million tokens and, per the announcement, more self-checking underway. You still sign it off.
xAI · Coding Agents · Agentic Engineering · Verification
REDAgentBench: not the stated guardrail, the verified one
1,661 executable attack cases against agent systems. In one diagnostic cohort, the agent had stated the rule beforehand in almost one in five violations.
Verification · Agentic Engineering · Automation
Mistral Regional Endpoints: the region becomes a setting, not a second provider
Mistral Regional Endpoints are generally available: inference in Europe or the US. The Priority Tier is public preview, GLM-5.2 is announced.
Mistral · Agentic Engineering · Verification
Muse Glimmer: a local judge becomes a realistic option
Meta releases Muse Glimmer: 30 billion parameters under Apache 2.0, built for local agents. The text states no scores. The role is the point.
Meta · Agentic Engineering · Verification · Automation
Session budgets in Claude Managed Agents: the cap is a pause, not a kill switch
Anthropic gives Managed Agents sessions a hard spend cap. A session that hits it pauses. Plus an advisor, an inference geo, and skills from GitHub.
Claude · Agentic Engineering · Automation · Verification
Codex CLI 0.147.0: not the plugins, the approval
Portable Agent Plugins, the opt-in MCP protocol of 28 July, and a flag called --approve-for-me. One line in the changelog moves the roles around.
OpenAI · Coding Agents · Agentic Engineering · Verification
Twelve CVEs across four agent frameworks: the boundary sits in the orchestration layer
Black Hat 2026: Check Point Research reports 12 CVEs across LangChain, Google’s ADK, Microsoft Agent Framework and CrewAI. Not in the model, in the plumbing.
Agentic Engineering · Verification · Automation
Claude Code on your own compute: execution moves back inside, the transcript does not
Claude Code sessions can run on your own infrastructure in public beta. Code and artifacts stay there, the conversation still goes to Anthropic.
Coding Agents · Agentic Engineering · Verification · Claude
Stateless MCP: state does not disappear, it changes layer
Google explains what the stateless core of the 28 July MCP specification means in production. A concrete migration item if you run your own servers.
Google · Agentic Engineering · Automation
Inference hooks: the veto sits before inference, the server sits with you
Anthropic holds every governed prompt on claude.ai, Cowork and Claude Code until your own server returns allow or deny. Beta for Claude Enterprise.
Claude · Verification · Agentic Engineering
Muse Code: not the new agent, the event log
Meta ships Muse Code with Muse Spark 1.2. The news is not the launch but the local event log and a plan that needs approval before work starts.
Coding Agents · Agentic Engineering · Verification · Meta
Claude Code 2.1.203 to 2.1.214: caps are an admission
200 subagents, 200 web searches, MCP calls backgrounded after two minutes, plus a row of closed permission bypasses. What that admits.
Claude · Coding Agents · Agentic Engineering · Verification
Kimi K3: open weights in the 3T class, and nothing verified yet
Moonshot AI introduces Kimi K3: 2.8 trillion parameters, 1M context, weights by 27 July. The coding numbers are vendor-reported.
Kimi · Coding Agents · Verification
Grok 4.5: not the benchmark rank, the cost per completed task
SpaceXAI introduces Grok 4.5: $2/$6 per 1M tokens and 4.2x fewer output tokens than Opus 4.8, at a lower SWE-Bench Pro score.
xAI · Coding Agents · Agentic Engineering · Automation
GPT-5.6 generally available: the table says less than the text above it
OpenAI makes Sol, Terra, and Luna generally available and publishes the full eval set. The coding lead in it is not consistent.
OpenAI · Coding Agents · Verification
PERFOPT-Bench: not the reported speedup, your own measurement
A benchmark measures coding agents on optimisation. No stack wins everywhere, and a 492.8x speedup shrinks to 13.1x once audited.
Coding Agents · Verification · Agentic Engineering · Automation
Managed agents in the background: convenient, and one more dependency
Managed agents in the Gemini API now run in the background and call remote MCP servers directly. What that saves you, and what control it costs.
Google · Agentic Engineering · Automation
Global workspace: the transcript is not the whole run
Anthropic’s J-lens reads internal states Claude never says out loud, among them detected bugs and prompt injections. What that is worth.
Claude · Verification · Agentic Engineering
Claude Code 2.1.199 to 2.1.201: the defaults move back toward the human
The default permission mode is now called Manual, questions no longer auto-continue, and subagents report API errors as errors. What that signals.
Claude · Coding Agents · Agentic Engineering · Verification
Mistral Leanstral 1.5: the model built for checking is the more interesting release
Mistral releases Leanstral 1.5 under Apache 2.0: Lean 4 proofs, 119B total, 6B active. Its agentic verification mode found 5 unreported bugs in 57 repos.
Mistral · Verification · Agentic Engineering · Coding Agents
Claude Code 2.1.198: browser agent GA, and the handoff moves into the draft PR
Claude in Chrome is generally available, and background agents open finished work as draft PRs instead of stopping to ask. What that means for your review.
Claude · Coding Agents · Agentic Engineering · Automation
DeepSeek V4: the API price is about to depend on the time of day
DeepSeek schedules V4 for mid-July: 1M-token context, peak hours billed at twice the rate, deepseek-chat and deepseek-reasoner retire on 24 July.
DeepSeek · Agentic Engineering · Automation
Fable 5 redeployed: the comeback is the smaller news
The US export controls are lifted, Fable 5 returns worldwide on July 1. More important than the comeback: the new severity framework for jailbreaks.
Agentic Engineering · Claude
Claude Sonnet 5: the price per token is not the price per task
Sonnet 5 gets close to Opus 4.8 at Sonnet prices. But the new tokenizer counts up to 1.35 times the tokens. Do the math per task, not per token.
Agentic Engineering · Claude · Coding Agents · Automation
Claude Code: an MCP layer that retries, and an agent panel you can read
OAuth and discovery retries for MCP, a readable agent panel, and about 37% less CPU while streaming. What the late-June updates actually harden.
Claude · Coding Agents · Agentic Engineering · Context Engineering
DeepSeek DSpark: speed from the inference layer, not from a new model
DeepSeek open-sources DSpark and the MIT stack DeepSpec: 60 to 85 percent faster generation on V4-Flash, lossless. No new model for you to re-evaluate.
DeepSeek · Agentic Engineering · Coding Agents
GPT-5.6 Sol: a preview is not yet production
OpenAI starts a limited preview of the GPT-5.6 series. Stronger at coding, narrower in access. What the data sheet says is not yet what holds in production.
OpenAI · Coding Agents · Verification
Agentic Engineering: not the model, the method
Which AI model you use keeps changing. What lasts is the method: structure, verification, architecture. My framework for building software with AI.
Agentic Engineering · Verification · Coding Agents
Delegate, don’t chat: Claude Tag moves into Slack
Anthropic’s Claude Tag: a shared, proactive agent in your Slack channel. Why it is more than a bot, and the one question it does not answer.
Agentic Engineering · Coding Agents · Automation · Claude
Claude Code hardens orchestration: reliability lives in the architecture
A depth limit for nested sub-agents, MCP calls that abort instead of hanging, context back after /clear. Three changelog entries, one theme: reliability.
Claude · Coding Agents · Agentic Engineering · Verification
Codex Remote is GA: the agent works, the human approves
OpenAI brings Codex Remote to general availability. The agent runs on your host, controlled from the app. Approval stays with the human.
OpenAI · Coding Agents · Automation · Verification
Agentic Engineering: not whether you use AI, but how
Not whether you use AI makes the difference, but how much structure and verification surrounds the output. My take after three technology shifts.
Agentic Engineering · Context Engineering · Verification
Kimi K2.7 Code: open weights lower the barrier, not the responsibility
Moonshot AI ships Kimi K2.7 Code: open weights, 256K context, low prices. Access gets broader, the checking stays your job.
Kimi · Coding Agents · Verification
Mistral connectors: keep access tight, make failures traceable
Mistral scopes connectors per API key, ships a step-by-step debugger, and brings them into the coding agent. Tight access, checkable failures.
Mistral · Coding Agents · Verification
Computer Use in Gemini 3.5 Flash: not the capability, the safeguards
Computer use becomes a native tool in Gemini 3.5 Flash. Browser, desktop and mobile control in one fast model. The real test is the safeguards.
Google · Agentic Engineering · Coding Agents · Automation · Verification
Mistral OCR 4: good because it shows its uncertainty
Mistral OCR 4 turns documents into structured data. The strong part is not the benchmark, but the per-word confidence. That is what makes checking possible.
Mistral · Automation · Verification
The Interactions API is GA: convenient, but who owns the loop?
Google’s Interactions API hits GA and becomes the default path for Gemini agents. Managed agents take over orchestration. The loop moves to the vendor.
Google · Agentic Engineering · Automation
Sakana Fugu: orchestration as a product, the verification stays with you
Sakana ships a multi-agent system as a single model behind an OpenAI-compatible API. Sakana takes over the orchestration, not the verification.
Agentic Engineering · Automation · Verification
Grok Build /goal: self-verification is built in, the proof is not
xAI ships /goal in Grok Build: an autonomous execute-and-verify loop. The loop now lives in the tool, the proof still stays your job.
xAI · Coding Agents · Automation · Verification
Fable 5 recalled: a model is not a foundation
Three days after launch, a US export-control directive cuts off Fable 5 worldwide. The real point is not the trigger, it is the dependence.
Agentic Engineering · Claude
Claude Fable 5: capability is not reliability
Anthropic introduces Fable 5 and Mythos 5, a jump beyond the Opus class. But whether software ships reliably is a different question from raw capability.
Agentic Engineering · Claude · Coding Agents · Verification
Loop engineering: build loops, not prompts
Stop prompting coding agents, design loops that prompt them. What loop engineering is, and why verification still stays with you.
Agentic Engineering · Coding Agents · Automation · Verification · Claude
NVIDIA Nemotron 3 Ultra: built for long agent runs, still yours to check
NVIDIA ships Nemotron 3 Ultra: 550 billion parameters, open weights, built for long agent runs. Cheaper and faster means more runs, and more to check.
NVIDIA · Agentic Engineering · Verification
Claude Opus 4.8: not the benchmarks, the review
Opus 4.8 is out. More than the benchmarks: the model lets code flaws slip through about four times less often. That is where the value is decided.
Agentic Engineering · Claude · Verification · Coding Agents · Automation
NVIDIA Nemotron 3 Nano Omni: perception gets cheap, accountability does not
NVIDIA ships Nemotron 3 Nano Omni: a small, open model for vision, audio, and text. It sees the screen and acts. Checking it stays your job.
NVIDIA · Agentic Engineering · Verification
GPT-5.5: more autonomy does not mean less checking
GPT-5.5 is stronger at agentic coding, at the same latency and with fewer tokens. The longer a model runs on its own, the more the checking matters.
OpenAI · Coding Agents · Verification