Research9 min read
Your Cyber Agent Does Not Learn Security From a Dataset
Why static SFT is necessary but insufficient, and how verified environments turn a language model into a governed defensive agent.
Research10 min read
Building the Experience Loop
How verified attempts become reusable evidence, training data, and regression tests for defensive agents.
Research8 min read
Measuring What the Model Actually Trains On
How to reason about the composition of cyber SFT data before measuring a post-training gain.
Research9 min read
Evaluating Reliability Beyond Pass@k
What to disclose when a benchmark result depends on the environment, verifier, and tool harness.