Divyansh Agrawal

I spend my time wandering between systems questions (inference, caching, distributed execution) and behavioral ones (evals, RL environments, agent decisions). I'm fascinated by how much the two ends inform each other.

Most of all, I care about learning how to evaluate these models well. The most interesting tests are the ones that expose unexpected failure modes, reward practical utility, and teach you something you didn't anticipate. I'm still early in figuring out my own research taste, especially when it comes to practical RL and alignment. I still don't know which long-term questions will matter most, but I know what feels worth digging into today.

Right now I'm working on FactorioRL, turning the game into an environment for RL policies and LLMs to interact with the same persistent world, so agents can acquire hierarchical skills, try to build factories, and figure out how to recover when their contraptions inevitably collapse.

Outside of this, I love movies, travelling, photography, random hackathons, and a problem with just enough shape to pull on. I'm always up for a chat about a good benchmark, a strange RL environment, infrastructure that works, or what you watched last weekend.

Topics I keep returning to

  • RL & agents
  • Evals & benchmarks
  • Infrastructure & inference
  • Model architecture
  • Quantisation
  • Alignment & reliability
  • Continual learning
  • Travelling
  • Photography

Recent work

  • EchoBench: a synthetic social web for evaluating epistemic arbitration.
  • FactorioRL: an agent-native Factorio environment for learning and testing control policies.
  • Hyperion: low-latency LLM routing, caching, and observability.