AI evaluator. Blind A/B evals of frontier coding agents. 1st place, Berkeley RDI AgentBeats (Software Testing). Building adversarial benchmarks.
- California
- https://joshuahickson.org/
Highlights
- Pro
Pinned Loading
-
LogoMesh/LogoMesh
LogoMesh/LogoMesh PublicLogoMesh is an agentic benchmark that grades AI-written code. It sends a coding task to an AI agent, then a panel of sub-agents judges the result. The final score tells you how much you can trust t…
-
-
solvent
solvent PublicA pattern language for loops that can pay their debts: vocabulary, a task-classification matrix, and pre-registered experiments on AI agent autonomy.
Python
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.



