Must Reads

ejholmes.github.io ·

This reframes the MCP rush around operability—debuggability, composability, and failure isolation—which is where agent integrations actually succeed or die in production.

06 The Map 10 The Law 15 The Gate

github.com ·

If AI contributes to your codebase, attaching the session to the commit is an audit primitive you’ll wish you had the first time you debug a regression or answer a compliance question.

13 The Documentation 16 The Validation 15 The Gate

bleepingcomputer.com ·

It updates the agent threat model with a concrete browser→localhost hijack path, which should immediately change how you expose local runtimes and authenticate tool endpoints.

14 The Immune System 15 The Gate 02 The Truth

tomtunguz.com ·

This is a practical playbook for the failure mode that quietly kills agent deployments—runaway tool loops—and it translates directly into middleware and runbook changes you can ship this week.

14 The Immune System 15 The Gate 06 The Map

theatlantic.com ·

If you build on Claude (or any closed provider), this is the clearest look at how a single high-stakes customer can turn “model policy” into your operational dependency overnight.

10 The Law 15 The Gate 16 The Validation

mksg.lu ·

Shows a concrete pattern where context compression and sandboxed tool outputs directly translate into longer-lived, more observable agent sessions.

06 The Map 07 The Tech Island 12 The Order

nanoclaw.dev ·

If your agents can touch a filesystem or network today, this is the clearest argument for assuming they’re malicious and designing isolation boundaries accordingly.

07 The Tech Island 10 The Law 15 The Gate

blogs.nvidia.com ·

Signals that agent orchestration is packaging into domain control planes for critical infrastructure—useful to copy even if you don’t work in telco.

09 The Orchestration 02 The Truth 08 The Artifacts

gist.github.com ·

A practical blueprint for turning specs into an agent-orchestrated verification loop—so humans approve outcomes instead of reading every token.

16 The Validation 14 The Immune System 06 The Map

cbsnews.com ·

If you build on frontier model providers, this is a live example of how ‘values’ and defense demand can collide into sudden gating risk you must architect around.

10 The Law 15 The Gate 16 The Validation

infoworld.com ·

It turns agent autonomy into unit economics you can actually design against, with guardrails that prevent loops and tool spam from destroying margins.

12 The Order 15 The Gate 09 The Orchestration

openai.com ·

It’s a concrete blueprint for moving from “tool-using chat” to durable, governable, long-horizon workflows where state—and therefore control—lives at the runtime boundary.

09 The Orchestration 06 The Map 15 The Gate

browser-use.com ·

If you’re running agents that execute code or browse the web, this shows an infrastructure pattern (micro-VMs + control plane + secretless) that meaningfully reduces blast radius without killing throughput.

07 The Tech Island 14 The Immune System 11 The Graph

machinelearning.apple.com ·

It’s rare, credible evidence of using LLM judgments to scale evaluation and improve ranking in a real production system—useful for anyone building model-assisted QA loops.

02 The Truth 16 The Validation

georgeguimaraes.com ·

It’s a pragmatic argument—and set of tactics—for catching the regressions that only appear when you test against real models, not mocked behavior.

16 The Validation 14 The Immune System

microsoft.com ·

One of the clearest blueprints for scaling “digital employees” by design—hierarchies, isolated subagents, and memory tiers you can map directly onto your own org’s task boundaries.

09 The Orchestration 07 The Tech Island 06 The Map

openai.com ·

Shows what it looks like when MCP stops being theory and becomes a bidirectional artifact loop—design ↔ code—with implications for how teams structure reviews and ownership.

11 The Graph 03 The Teamwork 09 The Orchestration

openai.com ·

A rare, outcome-tied benchmark story—15% NEPA drafting gains with a named eval (DraftNEPABench)—that you can use as a model for measuring agent ROI without hand-wavy demos.

16 The Validation 02 The Truth 06 The Map

anthropic.com ·

Treat this as the template for how to communicate (and operationalize) model change: what breaks, what behaviors emerge, and what governance signals users should plan around.

13 The Documentation 15 The Gate 08 The Artifacts

github.com ·

If you let agents run commands, this gives you a concrete, composable pattern for shrinking the blast radius (filesystem + network + command surface) without killing developer velocity.

07 The Tech Island 14 The Immune System 10 The Law

kanyilmaz.me ·

A concrete pattern for making tool access cheaper and more legible—freeing budget for the evaluations and audit trails that agent systems otherwise can’t afford.

06 The Map 12 The Order 13 The Documentation

geekwire.com ·

Desktop control is becoming a first-class runtime, and this acquisition signals that UI-level actuation (and its permission model) will be a core battleground for agent reliability.

09 The Orchestration 15 The Gate 03 The Teamwork

techmeme.com ·

If you’re shipping agents in production, scheduled execution is the moment you must formalize permissions, logging, and failure handling—because “unattended runs” are where incidents are born.

04 The Liberation 15 The Gate 16 The Validation

schneier.com ·

This is a crisp demonstration that your model’s reliability is downstream of data integrity, so you need provenance and validation controls before you trust any web-ingested ‘facts.’

02 The Truth 14 The Immune System 16 The Validation

openai.com ·

It shows what real adversaries actually do with models, which helps you design defenses around workflows, telemetry, and containment instead of prompt-level wishful thinking.

14 The Immune System 02 The Truth 15 The Gate

simonwillison.net ·

A tight, reproducible example of what ‘vibe coding’ actually yields in 45 minutes—and the hidden constraints (scope control, spec clarity, integration) you’ll need to operationalize.

04 The Liberation 06 The Map 03 The Teamwork

schneier.com ·

A crisp demonstration that web-sourced data pipelines are an integrity liability, forcing builders to treat provenance and dataset hygiene as first-class system design.

02 The Truth 14 The Immune System 16 The Validation

code.claude.com ·

A concrete blueprint for supervising a long-running local coding agent from web/mobile—exactly the kind of runtime/permission pattern that will define ‘desktop agents’ in production.

07 The Tech Island 15 The Gate 03 The Teamwork

openai.com ·

Case-study evidence of how attackers really blend AI with traditional ops—and the defensive signals/controls that matter when you’re building agent capabilities into products.

14 The Immune System 02 The Truth 15 The Gate

latent.space ·

Captures the key tactical shift in AI dev tools: the winning systems close review-to-fix loops, so you should redesign workflows around verification and iteration cadence—not just code generation.

03 The Teamwork 06 The Map 16 The Validation

vmfunc.re ·

If you ship agents into regulated workflows, this investigation is a warning shot: identity and verification vendors can quietly become surveillance infrastructure unless you audit the full pipeline.

10 The Law 16 The Validation 02 The Truth

github.com ·

Shows a practical pattern for parallelizing real engineering work with multiple coding agents in isolated worktrees—useful if you’re reorganizing dev around agent throughput instead of tickets.

09 The Orchestration 03 The Teamwork

bleepingcomputer.com ·

A concrete example of governance becoming product surface area: data boundaries are now enforced at the substrate level, not by user training or policy docs.

10 The Law 15 The Gate

normaltech.ai ·

This is the clearest attempt yet to turn “agent reliability” into an operational spec (dimensions + model comparisons + dashboard) that you can wire into your own release gates.

16 The Validation 14 The Immune System 02 The Truth

anthropic.com ·

Gives a concrete mental model for why assistants exhibit stable ‘personas,’ which directly affects how you design prompts, constraints, and red-team tests for long-running agents.

06 The Map 14 The Immune System

simonwillison.net ·

A practitioner’s reliability playbook for coding agents that turns ‘prompting’ into disciplined engineering practices you can institutionalize.

13 The Documentation 14 The Immune System 06 The Map

opper.ai ·

A simple, repeatable stress test that exposes brittle reasoning across frontier models—useful as a regression harness and a reality check for autonomy claims.

14 The Immune System 16 The Validation 02 The Truth

docker.com ·

Concrete, implementable patterns for isolating agent execution and handling secrets safely—exactly the failure mode teams hit when they operationalize ‘always-on’ agents.

07 The Tech Island 10 The Law 15 The Gate

openai.com ·

It’s a rare primary-source admission that a flagship coding benchmark is no longer measuring what everyone thinks—forcing you to redesign your eval strategy around contamination and provenance.

02 The Truth 16 The Validation 14 The Immune System

anthropic.com ·

Defines an actionable measurement layer for human–AI collaboration that can be tied to training, tooling changes, and outcome audits rather than vibes.

03 The Teamwork 16 The Validation 01 The Voyage

simonwillison.net ·

Clarifies that ‘Codex’ is a model-plus-harness product—and that the harness is where tool safety, UX, and reliability actually get decided.

06 The Map 11 The Graph 09 The Orchestration

boristane.com ·

One of the most copy-pastable operator patterns for preventing agent regressions: enforce intent artifacts (research/plan) as a gating layer before tool execution.

15 The Gate 13 The Documentation 01 The Voyage

garryslist.org ·

Hard usage data that quantifies where agents actually spend tool calls today—and therefore where the next non-dev vertical wedges are still open.

06 The Map 12 The Order 16 The Validation

shuru.run ·

Concrete local-first sandboxing approach (ephemeral microVMs + checkpoints) that maps directly to safer agent tool execution on developer machines.

07 The Tech Island 14 The Immune System 15 The Gate

georgeguimaraes.com ·

A clean reality check that durable execution is an architecture choice (state/workflows/queues), not a language choice—useful before you bet your agent platform on a runtime narrative.

09 The Orchestration 06 The Map 14 The Immune System

simonwillison.net ·

Defines an emerging ‘personal agent runtime’ abstraction (local, schedulable, containerized, message-driven) that changes how you design persistence, tooling, and UX for agents outside the cloud.

09 The Orchestration 07 The Tech Island 06 The Map

tinfoil.sh ·

If you’re shipping regulated or safety-critical agents, this is a concrete blueprint for turning “trust me, it’s the right model” into cryptographic evidence you can audit and contract around.

10 The Law 07 The Tech Island 02 The Truth

anuragk.com ·

Shows an extreme but increasingly relevant path for agent cost/latency budgets: model-specialized silicon that turns inference speed into a hardware property, not a software optimization.

07 The Tech Island 12 The Order 11 The Graph

june.kim ·

A rare, crisp articulation of runtime task-tree coordination (deps, parallelism, and human interrupts) that you can directly map onto production orchestrators and eval harnesses.

09 The Orchestration 11 The Graph

bleepingcomputer.com ·

A concrete data point that AI has already shifted attacker economics—use it to justify closing exposed admin surfaces and building rate-limit/identity controls before you add more agent automation.

14 The Immune System 02 The Truth

georgeguimaraes.com ·

A practical pattern for local semantic regression tests (embeddings + NLI) that helps prevent the quiet output drift that breaks agent workflows in month 2–3.

16 The Validation 14 The Immune System 02 The Truth

together.ai ·

A concrete new inference primitive (trajectory distillation + block-wise KV caching) that can move agent UX from ‘demo latency’ to ‘always-on’ economics without quality concessions.

12 The Order 16 The Validation 06 The Map

huggingface.co ·

If you’re betting on edge/on-prem agents, this marks the consolidation of the local inference stack into a durable platform, changing integration and lifecycle assumptions for llama.cpp/ggml users.

07 The Tech Island 11 The Graph 04 The Liberation

stripe.dev ·

This is rare production evidence of high-volume coding-agent throughput with explicit review checkpoints—useful for designing your own autonomy budgets, PR flows, and failure containment.

09 The Orchestration 03 The Teamwork 16 The Validation

anthropic.com ·

It’s the clearest blueprint this day for how to run a remediation-capable security agent with verification and human-review gates—exactly the loop most teams are about to attempt (and get wrong).

14 The Immune System 15 The Gate 16 The Validation

anthropic.com ·

First rigorous telemetry-backed measurement of how real users grant agent autonomy. Proves autonomy is rising and domain-specific risk patterns are emerging — validates proportional oversight (Principle 15).

16 The Validation 15 The Gate 14 The Immune System

lennysnewsletter.com ·

The head of Claude Code describes what happens after coding is 'solved' — the shift to collaborative AI products like Cowork that reshape professional workflows beyond just writing code.

03 The Teamwork 04 The Liberation

martinfowler.com ·

Fowler names the central tension of AI-assisted development: velocity without process discipline becomes a debt multiplier. Essential counterweight to the 'ship faster' narrative.

12 The Order 14 The Immune System 03 The Teamwork

simonwillison.net ·

Paul Ford's NYT piece (via Willison) captures the exact moment bespoke software costs collapsed. The clearest mainstream articulation of Principles 3-5 in action.

03 The Teamwork 04 The Liberation 05 The Joy

schneier.com ·

AI discovered 12 OpenSSL zero-days including a 9.8 CVSS critical. Concrete proof that AI security research is producing real results — and raising the stakes for defenders.

02 The Truth 14 The Immune System 15 The Gate

waymo.com ·

Waymo’s ‘advice, not control’ model is a concrete blueprint for human-in-the-loop gating that preserves scalability without surrendering safety authority.

15 The Gate 03 The Teamwork 01 The Voyage

schneier.com ·

It’s production-grade evidence that AI vulnerability discovery is now yielding serious zero-days—meaning your defensive posture must assume attackers get this too.

14 The Immune System 02 The Truth 15 The Gate

huggingface.co ·

If your enterprise agents fail in opaque ways, the trace-to-failure-signature approach here is the most actionable path to turning ‘agent weirdness’ into debuggable engineering work.

16 The Validation 02 The Truth 14 The Immune System

martinfowler.com ·

This reframes the current productivity boom as a debt accelerator and gives you the missing management primitive (risk tiering + supervision design) for not imploding at agent speed.

14 The Immune System 12 The Order 03 The Teamwork

openai.com ·

This is a rare, end-to-end benchmark that tests the full agent security loop (find→exploit→patch), giving you a concrete template for evals that match real autonomy risk.

16 The Validation 14 The Immune System 15 The Gate

schneier.com ·

It forces a re-think of “we encrypt traffic so we’re safe” by showing how inference metadata can leak sensitive intent—critical for any agent product with private tool calls.

14 The Immune System 10 The Law 16 The Validation

docker.com ·

Concrete, runnable guidance for a unified agent state store (vectors+graphs+docs+SQL) that lets you move from prompt stuffing to governed memory with reproducible deployment.

11 The Graph 06 The Map 07 The Tech Island

anthropic.com ·

A 1M-token context window plus consistency/coding gains changes the break-even point for prompt-vs-memory design, and you’ll want the primary-source details before rewriting your harness.

06 The Map 12 The Order

nist.gov ·

This is the clearest signal that agent interoperability and security controls are about to be standardized externally—meaning your internal agent architecture will need to map to real conformance targets, not vibes.

10 The Law 14 The Immune System 16 The Validation

machinelearning.apple.com ·

Per-output interactive proofs are a fundamentally different validation primitive than benchmarks—useful if you’re trying to make agents ‘provably right’ on high-stakes steps instead of merely ‘usually right.’

02 The Truth 16 The Validation

arxiv.org ·

Hard evidence that ‘helpful repo context files’ can backfire—this will change how you design coding-agent context boundaries and what you choose to standardize org-wide.

06 The Map 16 The Validation 02 The Truth

arxiv.org ·

Provides data you can operationalize: invest in curated skills libraries and stop expecting agents to bootstrap their own reliable skills from scratch.

02 The Truth 06 The Map 14 The Immune System

schneier.com ·

Reframes prompt injection as a multi-stage intrusion problem, giving security teams a shared mental model to design layered defenses instead of one-off prompt band-aids.

14 The Immune System 10 The Law 16 The Validation

kiankyars.github.io ·

One of the clearest ‘receipts’ that small multi-agent swarms can ship real systems code when you anchor them in tests and legible coordination—useful as a blueprint, not a demo.

09 The Orchestration 03 The Teamwork 06 The Map

machinelearning.apple.com ·

A rare, implementable pattern for cutting LLM cost/latency while preserving correctness—verified reuse is the kind of systems trick that actually scales agent workloads without silently corrupting outputs.

02 The Truth 14 The Immune System 12 The Order

techcrunch.com ·

A rare, concrete production-load datapoint that forces you to think in escalation rates, containment, and ops metrics—not agent demos.

04 The Liberation 09 The Orchestration 16 The Validation

simonwillison.net ·

Gives you the right failure model for agent adoption: the risk isn’t just buggy code, it’s losing the shared understanding needed to safely change anything later.

13 The Documentation 01 The Voyage

arxiv.org ·

One of the clearest primary-source blueprints for end-to-end autonomous research loops (generate→verify→revise) with evaluation evidence beyond leaderboard scores.

16 The Validation 03 The Teamwork 06 The Map

gist.github.com ·

If agents are going to operate on codebases long-term, treating SCM as a queryable temporal database is a plausible next substrate for retrieval, audit, and change control.

11 The Graph 06 The Map 13 The Documentation

washingtonpost.com ·

Signals where ‘agent features’ (voice, memory, personalization) turn into liability fast—consent and provenance can’t be bolted on later.

10 The Law 15 The Gate

simonwillison.net ·

Reframes ‘great engineer’ value in the agent era as direction-setting and coordination—useful for rewriting role expectations, interview loops, and on-call ownership.

01 The Voyage 03 The Teamwork 09 The Orchestration

simonwillison.net ·

A clear warning shot for engineering org design: AI boosts junior output immediately, but without explicit retraining loops you’ll create a mid-level capability cliff.

03 The Teamwork 04 The Liberation 13 The Documentation

tomtunguz.com ·

Explains the emerging default playbook for scaling agent capability—buy teams, not tools—which directly affects how you should plan build-vs-buy and retention for your own agent platform.

04 The Liberation 09 The Orchestration 12 The Order

github.com ·

A practical blueprint for moving agent capability to the edge—forcing you to rethink permissions, data flows, and failure modes when the phone (not the cloud) is the runtime.

07 The Tech Island 14 The Immune System 06 The Map

github.com ·

A masterclass in constraint-driven shipping: reading this will sharpen your instinct for what to simplify when models/tools bloat your agent harness.

08 The Artifacts 05 The Joy 12 The Order

huggingface.co ·

The clearest end-to-end example of ‘skills as a supply chain’: agents generating real CUDA, integrating into PyTorch, benchmarking, and publishing artifacts—plus the eval hooks you’ll need to trust it.

08 The Artifacts 16 The Validation 06 The Map

martinfowler.com ·

Puts language to the hidden cost center of agent adoption—supervisory task switching and lost shared context—so you can redesign roles and workflows before productivity collapses.

03 The Teamwork 13 The Documentation 09 The Orchestration

openai.com ·

A rare, concrete blueprint for scaling agent-heavy APIs where ‘who gets to do what, when’ is enforced by provably-correct real-time metering rather than ad-hoc rate limits.

10 The Law 16 The Validation 15 The Gate

openai.com ·

Shows the next practical step in assistant security: turning capability gating into explicit product modes/labels so enterprises can operationalize prompt-injection defenses.

15 The Gate 10 The Law 14 The Immune System

machinelearning.apple.com ·

If you’re building coding/program-synthesis agents, Cadmus is a turnkey way to run controlled, reproducible experiments without needing frontier-scale budgets—exactly what you need to debug failure modes.

02 The Truth 06 The Map 07 The Tech Island

blog.can.ac ·

It demonstrates, with a clean ablation, that the fastest way to ‘improve models’ in production is often to redesign the edit/tool harness—an immediately actionable lever for any coding or ops agent.

06 The Map 07 The Tech Island

simonwillison.net ·

First crisp case study of agentic supply-chain/reputation attack via GitHub + blogging, forcing builders to treat public write-access as a high-risk capability requiring gates and abuse response.

14 The Immune System 15 The Gate

waymo.com ·

A rare primary-source look at how a leading autonomy team argues safety validation for expanded fully driverless operations—useful as a benchmark for what ‘evidence’ should look like in your domain.

16 The Validation 02 The Truth 14 The Immune System

huggingface.co ·

This is concrete evidence of where tool-using agents fail in the wild (permissions, time, coordination) and gives you a template for building eval sandboxes that look like work, not benchmarks.

16 The Validation 06 The Map 07 The Tech Island

machinelearning.apple.com ·

A simple, deployable uncertainty signal you can add to reasoning models to drive gating, escalation, and monitoring without waiting for perfect calibration.

16 The Validation 14 The Immune System 02 The Truth

simonwillison.net ·

Turns tool distribution into an API primitive—inline, sandboxed Skills are a new supply-chain surface you’ll need to design for (portability, permissions, provenance).

07 The Tech Island 14 The Immune System 15 The Gate

nextgov.com ·

Shows an adoption playbook that doubles as a safety mechanism—voluntary waitlists and micro-training function as scalable gates for high-liability orgs.

03 The Teamwork 15 The Gate 12 The Order

schneier.com ·

A concrete multimodal attack example that forces you to extend prompt-injection thinking beyond text and into sensor/vision pipelines before you ship embodied or camera-enabled agents.

10 The Law 14 The Immune System 16 The Validation

openai.com ·

The clearest primary-source blueprint for how ‘agent-first’ teams win: shift engineering effort from writing code to building feedback-rich harnesses that make agents reliably productive.

01 The Voyage 06 The Map 07 The Tech Island

github.com ·

Practical, repo-level evidence that structured code indexing (tree-sitter) is becoming the winning pattern for context engineering in coding agents.

06 The Map 11 The Graph 16 The Validation

eric.lubow.org ·

A reminder with teeth: the differentiator is operations (reliability, scaling, incident discipline), which is exactly where agent deployments most often fail after the first successful demo.

14 The Immune System 16 The Validation 03 The Teamwork

simonwillison.net ·

Shows a concrete pattern for making agents trustworthy at scale: force them to ship executable demos/artifacts that a reviewer (or CI) can actually run, collapsing “agent said it works” into verifiable proof.

08 The Artifacts 14 The Immune System 16 The Validation

lennysnewsletter.com ·

Codifies lightweight team rituals and guardrails that prevent the most common AI product failure modes from hiding until launch—useful if you’re trying to institutionalize “minimum viable trust.”

01 The Voyage 14 The Immune System 15 The Gate

ainowinstitute.org ·

Gives you the external-facing playbook—right-to-information and legal design—that your internal logging/audit posture will be measured against when accountability becomes adversarial.

10 The Law 02 The Truth 16 The Validation

machinelearning.apple.com ·

If your bottleneck is multi-GPU inference, this is a rare architecture-level proposal that targets synchronization overhead directly—likely to change how you budget latency for agentic systems.

12 The Order 09 The Orchestration 07 The Tech Island

technologyreview.com ·

A useful corrective against ‘multi-agent’ hype that helps you recognize coordination theater and refocus on memory, goals, and reliability as the actual bottlenecks.

09 The Orchestration 06 The Map 16 The Validation

deepmind.google ·

One of the strongest primary-source attempts to tie frontier reasoning to concrete discovery workflows, offering a template for how to argue outcomes beyond generic capability claims.

02 The Truth 16 The Validation

huggingface.co ·

Offline, in-browser local inference changes your agent architecture, cost, and threat model; this preview shows the practical building blocks to make edge agents real.

07 The Tech Island 14 The Immune System 05 The Joy

tomtunguz.com ·

Frames the skills ecosystem as the new enterprise control plane—and warns why capability provisioning without provenance and permissions becomes a supply-chain security problem.

14 The Immune System 04 The Liberation 03 The Teamwork

interconnects.ai ·

A clear articulation of how coding-agent differentiation is shifting from benchmark wins to workflow usability—exactly the lens you need to choose and instrument agents in production.

16 The Validation 06 The Map 03 The Teamwork

nextgov.com ·

Signals where US policy is heading: federal preemption and clear regulatory lanes—meaning your compliance strategy can’t be state-by-state patchwork for long.

10 The Law 15 The Gate 12 The Order

microsoft.com ·

Shows how to turn evaluation into a durable system (datasets + leaderboard + real-device tests) that prevents regressions and makes progress legible in underrepresented settings.

16 The Validation 02 The Truth 15 The Gate

microsoft.com ·

A strong new mental model for making agent behavior learnable from few demonstrations by reducing action ambiguity via predictive structure—directly applicable to tool-using agents and robotics.

06 The Map 01 The Voyage 02 The Truth

openai.com ·

One of the cleanest ‘agent in the loop’ case studies where autonomy is judged by measurable cost reduction across massive experimental search—not by prompt quality.

02 The Truth 07 The Tech Island 16 The Validation

partnershiponai.org ·

A practical indicator of how large enterprises are standardizing governance playbooks (and what ‘responsible AI’ will look like in procurement and audit this year).

10 The Law 15 The Gate 13 The Documentation