skip to content

Blog

AI, in context.

The topic that occupies me most right now: AI and how it changes our work. Current releases and announcements, soberly assessed, with an eye on what actually matters in practice. The posts themselves are written by AI agents. I choose the topics, check the facts, and stand behind every text.

Period
Filter by topic

Qwen3.8-2.4T-A95B: the weights are open, the licence is not

5 min read

Qwen publishes Qwen3.8-2.4T-A95B: 2.4 trillion parameters, 95 billion active. The weights are there, the licence carries revenue thresholds.

Qwen · Coding Agents · Verification

Grok 4.6: self-checking during the run does not replace sign-off

4 min read

SpaceXAI introduces Grok 4.6 for long agent runs: $2/$6 per million tokens and, per the announcement, more self-checking underway. You still sign it off.

xAI · Coding Agents · Agentic Engineering · Verification

REDAgentBench: not the stated guardrail, the verified one

6 min read

1,661 executable attack cases against agent systems. In one diagnostic cohort, the agent had stated the rule beforehand in almost one in five violations.

Verification · Agentic Engineering · Automation

Mistral Regional Endpoints: the region becomes a setting, not a second provider

5 min read

Mistral Regional Endpoints are generally available: inference in Europe or the US. The Priority Tier is public preview, GLM-5.2 is announced.

Mistral · Agentic Engineering · Verification

Muse Glimmer: a local judge becomes a realistic option

6 min read

Meta releases Muse Glimmer: 30 billion parameters under Apache 2.0, built for local agents. The text states no scores. The role is the point.

Meta · Agentic Engineering · Verification · Automation

Session budgets in Claude Managed Agents: the cap is a pause, not a kill switch

5 min read

Anthropic gives Managed Agents sessions a hard spend cap. A session that hits it pauses. Plus an advisor, an inference geo, and skills from GitHub.

Claude · Agentic Engineering · Automation · Verification

Codex CLI 0.147.0: not the plugins, the approval

4 min read

Portable Agent Plugins, the opt-in MCP protocol of 28 July, and a flag called --approve-for-me. One line in the changelog moves the roles around.

OpenAI · Coding Agents · Agentic Engineering · Verification

Twelve CVEs across four agent frameworks: the boundary sits in the orchestration layer

5 min read

Black Hat 2026: Check Point Research reports 12 CVEs across LangChain, Google’s ADK, Microsoft Agent Framework and CrewAI. Not in the model, in the plumbing.

Agentic Engineering · Verification · Automation

Claude Code on your own compute: execution moves back inside, the transcript does not

5 min read

Claude Code sessions can run on your own infrastructure in public beta. Code and artifacts stay there, the conversation still goes to Anthropic.

Coding Agents · Agentic Engineering · Verification · Claude

Stateless MCP: state does not disappear, it changes layer

6 min read

Google explains what the stateless core of the 28 July MCP specification means in production. A concrete migration item if you run your own servers.

Google · Agentic Engineering · Automation

Inference hooks: the veto sits before inference, the server sits with you

5 min read

Anthropic holds every governed prompt on claude.ai, Cowork and Claude Code until your own server returns allow or deny. Beta for Claude Enterprise.

Claude · Verification · Agentic Engineering

Muse Code: not the new agent, the event log

5 min read

Meta ships Muse Code with Muse Spark 1.2. The news is not the launch but the local event log and a plan that needs approval before work starts.

Coding Agents · Agentic Engineering · Verification · Meta

Claude Code 2.1.203 to 2.1.214: caps are an admission

5 min read

200 subagents, 200 web searches, MCP calls backgrounded after two minutes, plus a row of closed permission bypasses. What that admits.

Claude · Coding Agents · Agentic Engineering · Verification

Kimi K3: open weights in the 3T class, and nothing verified yet

4 min read

Moonshot AI introduces Kimi K3: 2.8 trillion parameters, 1M context, weights by 27 July. The coding numbers are vendor-reported.

Kimi · Coding Agents · Verification

Grok 4.5: not the benchmark rank, the cost per completed task

4 min read

SpaceXAI introduces Grok 4.5: $2/$6 per 1M tokens and 4.2x fewer output tokens than Opus 4.8, at a lower SWE-Bench Pro score.

xAI · Coding Agents · Agentic Engineering · Automation

GPT-5.6 generally available: the table says less than the text above it

5 min read

OpenAI makes Sol, Terra, and Luna generally available and publishes the full eval set. The coding lead in it is not consistent.

OpenAI · Coding Agents · Verification

PERFOPT-Bench: not the reported speedup, your own measurement

4 min read

A benchmark measures coding agents on optimisation. No stack wins everywhere, and a 492.8x speedup shrinks to 13.1x once audited.

Coding Agents · Verification · Agentic Engineering · Automation

Managed agents in the background: convenient, and one more dependency

4 min read

Managed agents in the Gemini API now run in the background and call remote MCP servers directly. What that saves you, and what control it costs.

Google · Agentic Engineering · Automation

Global workspace: the transcript is not the whole run

4 min read

Anthropic’s J-lens reads internal states Claude never says out loud, among them detected bugs and prompt injections. What that is worth.

Claude · Verification · Agentic Engineering

Claude Code 2.1.199 to 2.1.201: the defaults move back toward the human

4 min read

The default permission mode is now called Manual, questions no longer auto-continue, and subagents report API errors as errors. What that signals.

Claude · Coding Agents · Agentic Engineering · Verification

Mistral Leanstral 1.5: the model built for checking is the more interesting release

4 min read

Mistral releases Leanstral 1.5 under Apache 2.0: Lean 4 proofs, 119B total, 6B active. Its agentic verification mode found 5 unreported bugs in 57 repos.

Mistral · Verification · Agentic Engineering · Coding Agents

Claude Code 2.1.198: browser agent GA, and the handoff moves into the draft PR

4 min read

Claude in Chrome is generally available, and background agents open finished work as draft PRs instead of stopping to ask. What that means for your review.

Claude · Coding Agents · Agentic Engineering · Automation

DeepSeek V4: the API price is about to depend on the time of day

3 min read

DeepSeek schedules V4 for mid-July: 1M-token context, peak hours billed at twice the rate, deepseek-chat and deepseek-reasoner retire on 24 July.

DeepSeek · Agentic Engineering · Automation

Fable 5 redeployed: the comeback is the smaller news

3 min read

The US export controls are lifted, Fable 5 returns worldwide on July 1. More important than the comeback: the new severity framework for jailbreaks.

Agentic Engineering · Claude

Claude Sonnet 5: the price per token is not the price per task

3 min read

Sonnet 5 gets close to Opus 4.8 at Sonnet prices. But the new tokenizer counts up to 1.35 times the tokens. Do the math per task, not per token.

Agentic Engineering · Claude · Coding Agents · Automation

Claude Code: an MCP layer that retries, and an agent panel you can read

4 min read

OAuth and discovery retries for MCP, a readable agent panel, and about 37% less CPU while streaming. What the late-June updates actually harden.

Claude · Coding Agents · Agentic Engineering · Context Engineering

DeepSeek DSpark: speed from the inference layer, not from a new model

3 min read

DeepSeek open-sources DSpark and the MIT stack DeepSpec: 60 to 85 percent faster generation on V4-Flash, lossless. No new model for you to re-evaluate.

DeepSeek · Agentic Engineering · Coding Agents

GPT-5.6 Sol: a preview is not yet production

4 min read

OpenAI starts a limited preview of the GPT-5.6 series. Stronger at coding, narrower in access. What the data sheet says is not yet what holds in production.

OpenAI · Coding Agents · Verification

Agentic Engineering: not the model, the method

6 min read

Which AI model you use keeps changing. What lasts is the method: structure, verification, architecture. My framework for building software with AI.

Agentic Engineering · Verification · Coding Agents

Delegate, don’t chat: Claude Tag moves into Slack

4 min read

Anthropic’s Claude Tag: a shared, proactive agent in your Slack channel. Why it is more than a bot, and the one question it does not answer.

Agentic Engineering · Coding Agents · Automation · Claude

Claude Code hardens orchestration: reliability lives in the architecture

4 min read

A depth limit for nested sub-agents, MCP calls that abort instead of hanging, context back after /clear. Three changelog entries, one theme: reliability.

Claude · Coding Agents · Agentic Engineering · Verification

Codex Remote is GA: the agent works, the human approves

4 min read

OpenAI brings Codex Remote to general availability. The agent runs on your host, controlled from the app. Approval stays with the human.

OpenAI · Coding Agents · Automation · Verification

Agentic Engineering: not whether you use AI, but how

4 min read

Not whether you use AI makes the difference, but how much structure and verification surrounds the output. My take after three technology shifts.

Agentic Engineering · Context Engineering · Verification

Kimi K2.7 Code: open weights lower the barrier, not the responsibility

3 min read

Moonshot AI ships Kimi K2.7 Code: open weights, 256K context, low prices. Access gets broader, the checking stays your job.

Kimi · Coding Agents · Verification

Mistral connectors: keep access tight, make failures traceable

4 min read

Mistral scopes connectors per API key, ships a step-by-step debugger, and brings them into the coding agent. Tight access, checkable failures.

Mistral · Coding Agents · Verification

Computer Use in Gemini 3.5 Flash: not the capability, the safeguards

4 min read

Computer use becomes a native tool in Gemini 3.5 Flash. Browser, desktop and mobile control in one fast model. The real test is the safeguards.

Google · Agentic Engineering · Coding Agents · Automation · Verification

Mistral OCR 4: good because it shows its uncertainty

3 min read

Mistral OCR 4 turns documents into structured data. The strong part is not the benchmark, but the per-word confidence. That is what makes checking possible.

Mistral · Automation · Verification

The Interactions API is GA: convenient, but who owns the loop?

4 min read

Google’s Interactions API hits GA and becomes the default path for Gemini agents. Managed agents take over orchestration. The loop moves to the vendor.

Google · Agentic Engineering · Automation

Sakana Fugu: orchestration as a product, the verification stays with you

5 min read

Sakana ships a multi-agent system as a single model behind an OpenAI-compatible API. Sakana takes over the orchestration, not the verification.

Agentic Engineering · Automation · Verification

Grok Build /goal: self-verification is built in, the proof is not

4 min read

xAI ships /goal in Grok Build: an autonomous execute-and-verify loop. The loop now lives in the tool, the proof still stays your job.

xAI · Coding Agents · Automation · Verification

Fable 5 recalled: a model is not a foundation

2 min read

Three days after launch, a US export-control directive cuts off Fable 5 worldwide. The real point is not the trigger, it is the dependence.

Agentic Engineering · Claude

Claude Fable 5: capability is not reliability

2 min read

Anthropic introduces Fable 5 and Mythos 5, a jump beyond the Opus class. But whether software ships reliably is a different question from raw capability.

Agentic Engineering · Claude · Coding Agents · Verification

Loop engineering: build loops, not prompts

4 min read

Stop prompting coding agents, design loops that prompt them. What loop engineering is, and why verification still stays with you.

Agentic Engineering · Coding Agents · Automation · Verification · Claude

NVIDIA Nemotron 3 Ultra: built for long agent runs, still yours to check

3 min read

NVIDIA ships Nemotron 3 Ultra: 550 billion parameters, open weights, built for long agent runs. Cheaper and faster means more runs, and more to check.

NVIDIA · Agentic Engineering · Verification

Claude Opus 4.8: not the benchmarks, the review

2 min read

Opus 4.8 is out. More than the benchmarks: the model lets code flaws slip through about four times less often. That is where the value is decided.

Agentic Engineering · Claude · Verification · Coding Agents · Automation

NVIDIA Nemotron 3 Nano Omni: perception gets cheap, accountability does not

3 min read

NVIDIA ships Nemotron 3 Nano Omni: a small, open model for vision, audio, and text. It sees the screen and acts. Checking it stays your job.

NVIDIA · Agentic Engineering · Verification

GPT-5.5: more autonomy does not mean less checking

3 min read

GPT-5.5 is stronger at agentic coding, at the same latency and with fewer tokens. The longer a model runs on its own, the more the checking matters.

OpenAI · Coding Agents · Verification