index / hughwalker08/pokemon_red_benchmark
pokemon_red_benchmark
Benchmarking Jev (TypeSafe's fast decision model) against Jev + a GPT-6 Sol planner at playing Pokémon Red, in a shared PyBoy harness built on NousResearch/pokemon-agent.
Research and EvalsControl LoopListedPython
Repository
Stars0
LanguagePython
License—
Created2026-09-27
Last pushed2026-09-27
Officialno
Judgment
Genuine Jev project0.78
gate 0.5
Uses Jev at runtime0.73
Is a meta list0.04
Is a reimplementation0.27
Quality scales
Substance0.99 / 3
gate 0.5
Docs0.00 / 3
Novelty2.01 / 3
Composite score0.35
Category distribution
Research and Evals0.96
Games and Simulation0.04
SDKs and Clients0.00
Applications0.00
Other0.00
Learning0.00
Integrations0.00
Agent and Dev Tooling0.00
Assigned categoryResearch and Evals
Category probability0.96
Category confidence0.96
PatternControl Loop
Pattern confidence0.88
Curation
StatusListed
Reasongate passed
README pickno
Overridenone
Sources
- search:jev in:name,description,readme created:2026-09-27
Provenance
Modeljev-1.13.0
Question setv2
Judged at2026-09-27 11:43 UTC