Agent Ops: RL trainers, loop engineering, and agent data & code layers

Show HN: I RL-trained an agent that trains models with RL (for –$1.3k). The agent learns to write and launch RL training jobs using inner-model performance as an outer-loop reward and open-sources code, weights, and infra — a concrete demo that agents can autonomously close model-development loops, which matters for agentic orchestration and reproducible artifacts (Principles 09, 08).

What is “loop engineering?”. The piece frames loop engineering as replacing manual prompts with persistent automated agent loops and shifts engineers toward designing, supervising, and debugging those workflows — a practical mindset change for building observable, maintainable agent systems (Principles 03, 06).

Adapter emerges from stealth with $17.8M to provide data infrastructure for AI agents. Adapter launches a data-infrastructure layer that gives agents controlled, auditable access to user data and raises $17.8M. That directly addresses the hard problem of safe, governed agent cognition — essential infrastructure if you want agents to act on sensitive systems and produce auditable outcomes (Principles 06, 11).

Concho AI turns enterprise codebases into a knowledge layer for AI agents. Concho converts sprawling codebases into a semantic knowledge layer that agents can query and act on. For outcome engineers this matters because it supplies contextual, executable knowledge to developer-facing agents, reducing brittle prompts and improving traceability of agent decisions (Principles 06, 11).

Solving 20 Erdős Problems with 20 Codex Accounts Running in Parallel. Star Fleet runs 20 parallel GPT-5.6 agents in isolated sandboxes to automate Lean 4 proofs and solve dozens of open problems. The project is a concrete example of large-scale agent orchestration, sandboxed execution, and producing verifiable artifacts — a template for building and governing high‑throughput agent fleets (Principles 09, 07).