Base camp · Tools

You've reached base camp.

Base camp is provisioned: two tools are on PyPI, and the leaderboard route is being pitched.

regexbench

Evaluate generated regular expressions: semantic equivalence, correctness, and ReDoS safety.

python apache-2.0 pypi 0.3.0

labloop

Agent-driven experiment loop: propose a change, run it time-boxed, keep it only if the metric improves.

python apache-2.0 pypi 0.2.0

regexleaderboard

Runs regexbench across models and publishes the numbers — scores, methodology, and a re-run command. No results yet; the table ships empty until there are.

evaluation planning
← Back to the lab Benchmarks GitHub