index / redhatpanda/jev-regress-bench
jev-regress-bench
Which of an agent's approved answers actually changed after a config edit? A reproducible bench: exact match vs embeddings vs a three-tier stack vs an LLM judge vs one Jev Choice.
Research and EvalsVerificationListedTypeScript
Repository
Stars0
LanguageTypeScript
LicenseMIT
Created2026-09-23
Last pushed2026-09-23
Officialno
Judgment
Genuine Jev project0.91
gate 0.5
Uses Jev at runtime0.85
Is a meta list0.04
Is a reimplementation0.07
Quality scales
Substance1.41 / 3
gate 0.5
Docs2.72 / 3
Novelty2.70 / 3
Composite score0.71
Category distribution
Research and Evals1.00
Other0.00
Applications0.00
Integrations0.00
Agent and Dev Tooling0.00
Games and Simulation0.00
Learning0.00
SDKs and Clients0.00
Assigned categoryResearch and Evals
Category probability1.00
Category confidence1.00
PatternVerification
Pattern confidence0.33
Curation
StatusListed
Reasongate passed
README pickno
Overridenone
Sources
- search:jev in:name,description,readme created:2026-09-23
- list:yibie/awesome-jev
Provenance
Modeljev-1.13.0
Question setv2
Judged at2026-09-24 11:55 UTC