Back to projects
AI / ML

Hotpath - AI Performance Agent

An AI agent that makes code faster and proves it - every change must pass the locked tests and beat the measured noise floor.

Hotpath - AI Performance Agent

Overview

Most AI coding tools will happily tell you a change is faster. None of them can show you why you should believe it. Hotpath is built the other way round: the model is allowed to be creative and wrong, and a harness it cannot reach decides what is true. Point it at a repository and one command runs nine stages - clone, work out how to test and benchmark it, confirm the baseline is green and not flaky, search for optimizations, open a draft pull request with one commit per verified change, then watch the repository's own CI. A planner model reads the profile and proposes hypotheses; worker models write each one as a patch in its own isolated git worktree. Then the deterministic half takes over: the locked test suite runs first, and only code that passes is ever timed. A change is kept only if it is correct and faster than the noise. The acceptance bar is max(min_speedup, 1 + noise_multiplier x measured_noise), and the 95% bootstrap confidence interval over 2,000 resamples must exclude 1.0. A noisy benchmark makes Hotpath stricter, never more permissive. Released on PyPI as hotpath-agent.

Highlights

  • Opened a real draft pull request against jaraco/inflect - a repository nobody on the team wrote - from a bare GitHub URL in 5 minutes 3 seconds with no human step in between: 214 tests green, 8 candidates, 2 accepted, 1.472x faster.
  • Locked-path enforcement happens in code before a single byte is written, so a patch that tries to edit the tests it must pass is rejected before it runs.
  • Noise-aware accept rule: a 2,000-resample bootstrap CI must exclude 1.0. On one run a change measured 1.004x faster and was thrown out as noise; on another the benchmark's own 13% noise raised the bar to roughly 26% rather than lowering the standard.
  • Refuses rather than inventing a result. On mahmoud/boltons it reported 7 candidates, 0 accepted and opened no pull request; on life4/textdistance it stopped at the baseline because the suite's 200ms property-test deadline fails under load.
  • Candidate code runs in a fail-closed Docker sandbox - no network, read-only source, non-root, resource quotas - with an explicit opt-in local backend for trusted repositories.
  • On an H100, took a transformer from 965 to 1,409 tokens/sec (1.46x). Of 51 attempts, 22 broke correctness and 18 were not faster than noise; 3 shipped.
  • Ships a live dashboard, a generated-and-validated benchmark when a repo has none, leave-one-out ablation, and a preflight that prices a run before you spend anything.

Built with

PythonOpenAI APIDockerGitGitHub ActionsVercel
Chat with My AI Twin