Agent safety, open models, and cost: 5 signals for outcome engineers
Quoting Thibault Sottiaux reveals Codex in full-access mode can delete user files when sandboxing and auto-review protections are disabled. This exposes urgent gaps in agent sandboxing and governance — if your agents get file-system access, enforce sandbox defaults, audit auto-review toggles, and treat destructive capabilities as deploy-time kill switches (Principles 10 & 14).
LM Studio expands beyond chat with Bionic, a new AI agent app for open models launches Bionic, a Mac agentic app that runs open models locally, supports coding and document workflows, and offers secure cloud scaling. This lowers the barrier to building agentic workflows on-device and forces outcome engineers to design for mixed local/cloud execution, secure tool integration, and role-specific model selection (Principles 09 & 07).
Researcher poisons open-weight AI model for under $100 shows a backdoor implanted in an open-weight model in under an hour for under $100, proving model-poisoning is cheap and practical. Outcome engineering must treat open models as hostile inputs: add provenance checks, runtime behavior monitoring, and revocation/rollback paths to your agent pipelines to avoid silent supply-chain compromises (Principles 06 & 14).
The State of Open Source AI — V1.0 (July 2026) reports open-weight models now dominate production tokens and that capability gaps are closing as inference costs collapse, shifting value to agentic orchestration. That signal means outcome engineers should invest less in model-selection drama and more in orchestration layers, context management, and cost-aware routing across models and tools (Principles 09 & 11).
NVIDIA Vera Rubin Maximizes Intelligence per Dollar for Post-Training Workloads — a Key Metric for Agentic AI introduces a focus on intelligence-per-dollar for continuous post-training and RL-driven agent learning. Use this to prioritize post-training orchestration, compare agent learning approaches by cost-normalized gains, and bake intelligence-per-dollar into your outcome scorecards and deployment gates (Principles 12 & 09).