Roseram NewsAll stories

Developer, Coding & Automation · Microsoft Research · October 7

Agent Lightning trains coding agents inside their real working harness

Microsoft’s rebuilt reinforcement-learning framework connects training to the tools agents actually use.

Tools: Agent Lightning · Qwen

What happened

Microsoft Research introduced Agent Lightning v1.0 on October 7. Its approach includes the deployment harness directly in reinforcement learning instead of recreating the agent inside a separate training framework. The control plane is roughly 3,500 lines and supports Kubernetes execution. Researchers report raising Qwen3.5-9B from 41.8% to 56.4% Pass@1 on SWE-bench Verified with about 6,000 training samples.

Roseram analysis

An agent is more than its model: tool behavior and execution context affect outcomes. Training in the real harness tackles that mismatch. The reported gain describes one setup, not every model or repository. Teams evaluating the framework should record dataset provenance, compute costs, tool permissions, and results on held-out work; an open-source control plane does not make training free.

Read the underlying record.

  1. 01Microsoft Research announcement ↗