Modularium Research

AI research for model behavior we can understand.

We study how models act under pressure and turn the results into clear evidence.

Current study

What happens when the test pushes back?

We build evaluations where models use tools, make plans, and face constraints.

Evaluation track

Latest work

Small studies, clear claims.

Notes

What we are thinking through.

  1. Evaluation should explain behavior, not just rank systems
  2. Agent audits need pressure, not theater
  3. Mechanistic evidence as a publication standard
  4. Small surface area, high scrutiny