As the first quality hire, you will build and own the evaluation infrastructure for autonomous agents operating in national security environments. You’ll develop simulation harnesses for long-horizon behaviors, create labeled datasets for training and regression, and make critical ship/no-ship decisions for on-premise deployments that directly protect people and confront adversaries.
Evals Lead at well-funded national security AI startup
Join a well-funded AI startup as their first Evals Lead and take ownership of the quality function for autonomous agents deployed in high-stakes national security missions. You’ll build the simulation harnesses and labeled datasets that govern how agents plan and act over weeks-long operations, working alongside a founding team from the CIA, FBI, and top defense-tech firms. This isn't just about measurement; you'll own the ship/no-ship decisions for on-premise deployments that confront real-world adversaries. If you're ready to define the evaluation discipline for agents where per-episode metrics simply don't cut it, this is the role for you.
Want to apply for this role?
Jack finds you jobs at companies like Confidential company. Talk to Jack to get considered for roles that fit what you're great at.
Location
Washington, D.C., United States
Compensation
$250k-$280k + Equity
Company
Confidential company
Role overview
About the company
Well-funded national security AI startup
What you will do
- Operate the agent platform daily to identify failure modes and build automated tooling to catch regressions in agent planning and decision-making.
- Develop simulation harnesses that condense weeks of autonomous agent activity into overnight tests, ensuring reliability for fixed-cadence on-premise deployments.
- Create and manage versioned, labeled datasets used for both regression testing and training or reward modeling for complex agentic workflows.
Who this is a fit for
- Proven experience owning evaluation suites for shipped LLM or agent products, with a deep understanding of when LLM-as-judge succeeds or fails.
- Strong Python skills and familiarity with agent harness internals such as memory, context assembly, tool selection, and process supervision.
- Background in building labeled datasets for model training and experience with production evaluation frameworks like Promptfoo, Braintrust, or LangSmith.
Why this role is remarkable
- Lead the evaluation discipline for autonomous agents that execute weeks-long missions against real-world adversaries in high-stakes environments.
- Well-funded by top-tier VCs and led by a founding team with deep experience from elite intelligence agencies and leading defense-tech companies.
- Full ownership of the evaluation function from day one, including founding-team equity and a direct impact on operational mission success.
How Jack & Jill work together
Jack gets to know what you're great at and what you want next, then searches 15 million jobs daily and helps you discover roles at companies like this.
Meet Jack
What happens next?
Jack’s an AI agent for job searching and career coaching. He works for you.
Jill is the AI recruiter working for the company. She recruits from Jack’s network.
If your profile’s a match and Confidential company wants to meet, Jill will make the intro. In the meantime, Jack will send you excellent alternatives.