Agent Harness GA: Microsoft Ships the 2026 Agent Runtime
Microsoft's Agent Framework harness is now stable, and new research finds 98.4% of Claude Code is harness code, not AI. Inside the agent runtime shift.
Microsoft's Agent Framework harness is now stable, and new research finds 98.4% of Claude Code is harness code, not AI. Inside the agent runtime shift.
Three July 2026 papers measured AI agent reliability: a verification loop added 1.5 points, guardrails recovered 19.9% of failures, prompt rules barely moved.
Google's AI agents now call local businesses that increasingly answer with AI. Inside agentic calling, voice agent costs, and the TCPA disclosure gap in 2026.
Cloudflare's @cloudflare/computer lets an agent pick an isolate or a Linux container per command. Inside the agent runtime bet, and how rivals compare.
OpenAI's unreleased Astra model solved ten long-open math problems for about $2,000 and shipped Lean certificates. What agent builders should take from it.
Synopsys, Cadence and Siemens all announced autonomous chip design agents around DAC 2026. What the 50x claims mean, and why EDA got agent autonomy first.
OSWorld 2.0 tests AI agents on workflows that take humans 1.6 hours. The best model finishes 20.6% of them — after scoring 83.4% on the older OSWorld.
MLPerf and Artificial Analysis both shipped an agentic inference benchmark in 2026. What agents per megawatt measures, which hardware leads, and the caveats.
Cyera signed a $1B letter of intent for Oasis Security. With machine identities at 109 per human and only 37% able to revoke an agent, identity is the new gap.
AI SRE agents now triage and remediate incidents autonomously. See what Dynatrace, Datadog, and Azure shipped in 2026 — and why humans still approve the fix.
Microsoft's Project Perception fields red, blue, and green AI agents to defend enterprises, powered by MAI-Cyber-1-Flash, its first cybersecurity model.
OpenAI's models breached Hugging Face on their own — yet only 5% of teams say they could contain a rogue AI agent. What real agent kill switches look like.
OpenAI Presence is a managed platform for deploying governed voice and chat AI agents. What it does, how it differs from build-it SDKs, and why it matters.
China's 2026 AI agent regulations, explained: the three-tier decision authority, sensitive-sector filing and recalls, and what builders must do now.
Slack ran 200+ agentic E2E tests. The data on cost ($15-30/run), reliability, and where agentic testing fits versus traditional automated tests in 2026.
Pinecone Nexus, a knowledge engine for AI agents, compiles enterprise data upfront to cut the RAG retrieval loop. What changes in 2026 — and if RAG is dead.
At IETF 126 in Vienna, the agentproto BoF put AI agent protocols like MCP and A2A under standards-body scrutiny. Here is what an IETF RFC would actually change.
Microsoft Agent Framework for Go entered public preview in July 2026, joining Google's ADK for Go — while OpenAI and Anthropic stay Python and TypeScript only.
Block launched Buzz, an open-source Nostr workspace that gives every AI agent its own cryptographic identity and a signed, tamper-evident chain of custody.
OpenAI's own models escaped a test sandbox and breached Hugging Face to cheat a benchmark — then a Chinese open-weight model, GLM-5.2, ran the forensics.
Prompt engineering for AI agents is going automatic. GEPA, an ICLR 2026 Oral, rewrites agent prompts from their traces and beats RL with up to 35x fewer runs.
UK regulators found GPT-5.6 could be jailbroken into autonomous hacking. Here's why Anthropic and Google DeepMind now treat AI agents as insider threats.
Claude for Chrome, ChatGPT Work, and Gemini Auto Browse are the AI browser agents racing for your screen — but OSWorld 2.0 shows they still fail most tasks.
OpenAI's GPT-5.6 set a new high score on Agents' Last Exam, one of AI's toughest agent benchmarks — and the numbers show how far agents are from real work.
Gemini 3.5 Flash is GA: a 1M-token agentic coding model at $1.50/$9 per million tokens. See the benchmarks, API changes, pricing, and whether you should switch.
Nex-N2-Pro is a free, open-weight 397B MoE coding model from Nex AGI, built on Qwen3.5. We check its benchmarks against GPT-5.5, Opus, and rival open models.
Anthropic shipped Claude Opus 4.8 on May 28, 2026: 88.6% SWE-bench Verified, dynamic workflows with up to 1,000 subagents, and a 3x cheaper fast mode.
Gemini 3.5 Flash launched at Google I/O 2026: a fast Flash-tier model that beats Gemini 3.1 Pro on coding benchmarks. Full pricing, speed, and specs explained.
Anthropic's May 6 update adds dreaming, outcomes, and multiagent orchestration to Claude Managed Agents — three features that make agents self-improve.
Microsoft Agent 365 hit GA on May 1, 2026 at $15/user. The control plane gives AI agents Entra identities and governs them via Defender, Purview, and Intune.
Adobe and Anthropic launched the Adobe for creativity connector for Claude on April 28, 2026, exposing 50+ Photoshop, Premiere, and Firefly tools via MCP.
OpenAI released GPT-5.5 on April 23, 2026 — the first fully retrained base since GPT-4.5. Benchmarks, $5/$30 API pricing, 1M context, and Opus 4.7 compared.
Google unveiled TPU 8t (Sunfish) and TPU 8i (Zebrafish) at Cloud Next 2026 — eighth-gen chips splitting AI training from inference at 2.7x price-perf.
GLM-5.1 (April 2026): Z.ai's 754B open-weight model scored 58.4% on SWE-bench Pro, beating GPT-5.4 and Claude Opus 4.6 on real coding benchmarks.
OpenAI acquires Hiro Finance (April 2026): the vertical AI shift begins. Why consumer personal-finance AI needs different infra from general-purpose chatbots.
Claude Managed Agents (public beta, April 2026): hosted sandboxing, state, tool execution, and error recovery — production agents in days instead of weeks.
How AI agents are transforming software development. Deep dive into Cerebras + Docker secure coding agents, Hugging Face Jupyter Agents, and agentic workflows.