Skip to main content

Evals Lead at well-funded national security AI startup

Join a well-funded AI startup as their first Evals Lead and take ownership of the quality function for autonomous agents deployed in high-stakes national security missions. You’ll build the simulation harnesses and labeled datasets that govern how agents plan and act over weeks-long operations, working alongside a founding team from the CIA, FBI, and top defense-tech firms. This isn't just about measurement; you'll own the ship/no-ship decisions for on-premise deployments that confront real-world adversaries. If you're ready to define the evaluation discipline for agents where per-episode metrics simply don't cut it, this is the role for you.

Want to apply for this role?

C

Jack finds you jobs at companies like Confidential company. Talk to Jack to get considered for roles that fit what you're great at.

Location

Washington, D.C., United States

Compensation

$250k-$280k + Equity

Company

Confidential company

Talk to Jack

Role overview

As the first quality hire, you will build and own the evaluation infrastructure for autonomous agents operating in national security environments. You’ll develop simulation harnesses for long-horizon behaviors, create labeled datasets for training and regression, and make critical ship/no-ship decisions for on-premise deployments that directly protect people and confront adversaries.

About the company

Well-funded national security AI startup

What you will do

  • Operate the agent platform daily to identify failure modes and build automated tooling to catch regressions in agent planning and decision-making.
  • Develop simulation harnesses that condense weeks of autonomous agent activity into overnight tests, ensuring reliability for fixed-cadence on-premise deployments.
  • Create and manage versioned, labeled datasets used for both regression testing and training or reward modeling for complex agentic workflows.

Who this is a fit for

  • Proven experience owning evaluation suites for shipped LLM or agent products, with a deep understanding of when LLM-as-judge succeeds or fails.
  • Strong Python skills and familiarity with agent harness internals such as memory, context assembly, tool selection, and process supervision.
  • Background in building labeled datasets for model training and experience with production evaluation frameworks like Promptfoo, Braintrust, or LangSmith.

Why this role is remarkable

  • Lead the evaluation discipline for autonomous agents that execute weeks-long missions against real-world adversaries in high-stakes environments.
  • Well-funded by top-tier VCs and led by a founding team with deep experience from elite intelligence agencies and leading defense-tech companies.
  • Full ownership of the evaluation function from day one, including founding-team equity and a direct impact on operational mission success.

How Jack & Jill work together

Jack
I get to know what you’re great at, then find roles you’d never find yourself.
Jill
I recruit from Jack’s network and make the intro when I spot a great match.
Thumbnail for Meet Jack

Jack gets to know what you're great at and what you want next, then searches 15 million jobs daily and helps you discover roles at companies like this.

Meet Jack

What happens next?

Jack’s an AI agent for job searching and career coaching. He works for you.

Jill is the AI recruiter working for the company. She recruits from Jack’s network.

If your profile’s a match and Confidential company wants to meet, Jill will make the intro. In the meantime, Jack will send you excellent alternatives.

Learn about Jack

Ready to find your next role?

Talk to Jack for 10 minutes and see your first matches.