# High-assurance software, proven by evidence

At Code Metal Research, we advance AI and formal methods to make high-assurance software possible for systems of consequence. We work across AI, programming languages, and formal verification, with a focus on turning research into real systems where correctness, performance, and trust matter.

## What We Work On

01

### AI Generated Systems
Making AI-generated and autonomous software trustworthy.

02

### AI-Assisted Formal Methods
Leveraging the power of AI to advance, accelerate, and automate theorem proving, verification, program synthesis, and proof automation.

03

### High-Assurance Systems of Consequence
Applying formal methods to defense, cyber, critical infrastructure, and domains where software failure is not an option.

## How We Work

### Apply research to real systems
Our ideas are grounded in real engineering problems, real codebases, and systems where correctness, performance, and trust matter. We move beyond theory by applying research where the consequences of software failure are real.

### Publish work that moves the field
We contribute to the research community by sharing findings, benchmarks, technical reports, and open problems that advance the state of the field. Our researchers publish, present, and collaborate across papers, conferences, seminars, and academic partnerships.

### Learn and lead from the frontier
We bring together researchers, engineers, and domain experts across AI, programming languages, systems, and formal methods to challenge assumptions and advance what high-assurance software can become.

## Formally Speaking

Formally Speaking is our seminar series for researchers working at the seam between formal methods and AI. We invite people whose work we are learning from to share what they are building, what they are still figuring out, and what the field needs to solve next.

- Jun 30  
**Nada Amin**  
Harvard · MidSpiral  
*Formal Verification + AI: MidSpiral's Practical Approach*

- Jul 7  
**Aws Albarghouthi**  
UW–Madison · AWS  
*Verifying, Heavy & Light*

- Jul 13  
**Ilya Sergey**  
NUS  
*Using Lean as a Multi-Modal Meta-Verifier*

- Jul 21  
**Xinyu Wang**  
U-M  
*Superoptimization for Database Queries*

- Jul 28  
**Ranjit Jhala**  
UCSD  
*Flux: Refinement Types for Verified Rust Systems*

- Aug 11  
**Adam Chlipala**  
MIT  
*Scaling Formal Verification to Complete Hardware-Software Stacks*

- Aug 18  
**John Regehr**  
Utah  
*Translation Validation for LLVM's AArch64 and RISC-V Backends*

## Research Papers & Articles

**The Trust Problem Has Shifted: What Formal Verification Can and Cannot Guarantee About AI-Generated Code**  
A clear-eyed technical assessment of formal verification for AI-generated code: which approaches are credible, what barriers remain, and where the market will emerge first.  
*July 8, 2026*  
[Read more →](/content/research/the-trust-problem-has-shifted-what-formal-verification-can-and-cannot-guarantee-about-ai-generated-code/index.html)

**The Real Cost of Leaving NVIDIA**  
What Automated Transpilation Actually Costs, and What It Doesn't  
*June 4, 2026*  
[Read more →](/content/research/the-real-cost-of-leaving-nvidia/index.html)

**AI-generated code that works — and proves it**  
How Code Metal combines AI with formal methods to build trusted code translation systems, and welcoming Prof. Loris D'Antoni as our first Code Metal Scholar.  
*May 18, 2026*  
[Read more →](/content/research/ai-generated-code-that-works-and-proves-it/index.html)

**Counting Without Running: Evaluating LLMs' Reasoning About Code Complexity**  
Introduces gpuFLOPBench, a benchmark containing 577 CUDA kernels to evaluate whether language models can predict floating-point operation counts without execution, revealing limitations in understanding hardware-specific performance details.  
*December 4, 2025*  
[Read more →](/content/research/counting-without-running-evaluating-llms-reasoning-about-code-complexity/index.html)

## Code Metal Research in the Community

Paper

### AAMAS Conference 2026
Senior AI Researcher Sanjna Ravichandar presented on applying reinforcement learning to optimize logistics for critical national infrastructure.

Keynote

### IEEE Conference 2026
Principal Research Scientist Dr. Niranjan Hasabnis delivered a keynote at the IEEE Annual Computing and Communication Workshop.

Paper

### NeurIPS 2025
Researchers Ellie Kitanidis and Cole Hunter presented at NeurIPS on code representations and the limitations of today's code embeddings.

Keynote

### ITP Conference 2025
Principal Research Scientist Dr. Laura Titolo delivered a keynote talk at the 16th International Conference on Interactive Theorem Proving.

See us next at:

- **ACL 2026**: ParaCodex: A Profiling-Guided Autonomous Coding Agent for Reliable Parallel Code Generation and Translation  
July 2–7, 2026
- **NSAD 2026**: Numerical & Symbolic Abstract Domains  
October 3–9, 2026
- **Dagstuhl Seminar**: The Next 20 Years of Computer-Assisted Theorem Proving  
February 21–26, 2027

## Research Leadership

**Dr. Ellie Kitanidis**  
AI Research Lead  
*Past Experience: OpenAI | UC Berkeley | Stanford*

**Dr. Laura Titolo**  
Formal Methods Research Lead  
*Past Experience: NASA Langley Research Center | National Institute of Aerospace*

**Dr. Loris D'Antoni**  
Code Metal Scholar · Professor at UCSD  
*Past Experience: AWS | UW-Madison | University of Pennsylvania*

We have a growing team of researchers from leading educational institutions and applied industry organizations, including Stanford, Harvard, MIT, Cornell, UC Berkeley, Google, Intel, Bloomberg, and AWS.

## Advance the frontier with us

We're growing the team with researchers, engineers, and builders who want to advance provable AI and bring it into real systems. If that sounds like work you want to help shape, explore our open roles and join us.
