Jev-MedQA
Experiments running Jev on medical multiple-choice question answering.

- GitHub stars
- 1 as of 2026-09-25
- Forks
- 0
- Language
- Python
- License
- MIT
- Created
- 2026-09-25
- Last updated on GitHub
- 2026-09-25
- Added to AionEdge
- 2026-09-25

What Jev does in Jev-MedQA
The project describes itself this way: Experiments running Jev on medical multiple-choice question answering. In the Benchmarks & research category, Jev typically decides labels and probabilities on fixed test sets, so accuracy, calibration and latency can be measured.
Verification
Verified. We confirmed this repository through the GitHub API and located the code that calls Jev at a pinned commit: view the Jev code ↗. Upstream code changes over time, so check that the file still exists at the current commit before citing it.
How Benchmarks & research projects use Jev
Benchmarks and research projects test Jev rather than just use it. They compare it with language models and other classifiers, check whether its probabilities are well calibrated, and explore where the decision-model approach breaks down.
Related Jev projects in Benchmarks & research
- decision-benchHanno-Labs/decision-benchOpen benchmark runtime for document-grounded decision models.★ 1
Python - jev-synergy-screeningPistachioAIHQ/jev-synergy-screeningA Jev title-and-abstract screening experiment compared with Cohen Abstract Triage labels for an ADHD review.★ 1
Python
jev-agent-failure-benchmarkTokenTrim/jev-agent-failure-benchmarkA benchmark using Jev to attribute multi-Agent failures to an Agent, step and error type.★ 2
Python- jev-calibration-auditjujumilk3/jev-calibration-auditAudits Jev calibration, option-wording effects and Korean judgments through public APIs and datasets.★ 0
Python - jev-transcript-screenerkeduseworku/jev-transcript-screenerDemo and evaluation harness that uses Jev to rank solver transcripts for review by a more costly judge.★ 0
Python - jev-vs-clefrmax-ai/jev-vs-clefLaunch-window comparison harness for Jev and Cloudflare Clef decision models on synthetic decision tasks.★ 0
Python