aionedge.
Jev project · Benchmarks & research

jev-transcript-screener

keduseworku/jev-transcript-screener

Demo and evaluation harness that uses Jev to rank solver transcripts for review by a more costly judge.

View on GitHub ↗See the Jev code ↗
GitHub stars
0 as of 2026-10-05
Forks
0
Language
Python
License
MIT
Created
2026-09-24
Last updated on GitHub
2026-10-05
Added to AionEdge
2026-10-05
Built by this developer
keduseworku

What Jev does in jev-transcript-screener

Asks literal questions about redacted excerpts and combines probabilities into screening scores. The bundled corpus is synthetic; demo results do not establish generalization to real solver traffic. Catalog verification inspected README and source; the application was not independently runtime-tested.

Verification

Verified. We confirmed this repository through the GitHub API and located the code that calls Jev at a pinned commit: view the Jev code ↗. Upstream code changes over time, so check that the file still exists at the current commit before citing it.

How Benchmarks & research projects use Jev

Benchmarks and research projects test Jev rather than just use it. They compare it with language models and other classifiers, check whether its probabilities are well calibrated, and explore where the decision-model approach breaks down.

Related Jev projects in Benchmarks & research

  1. jev-calibration-auditjujumilk3/jev-calibration-auditAudits Jev calibration, option-wording effects and Korean judgments through public APIs and datasets.★ 0
    Python
  2. jev-vs-clefrmax-ai/jev-vs-clefLaunch-window comparison harness for Jev and Cloudflare Clef decision models on synthetic decision tasks.★ 0
    Python
  3. jev-comparewestsmith-open/jev-compareSide-by-side Jev and Gemini evaluation harness comparing typed decisions, requests, latency, cost and calibration.★ 0
    HTML
  4. decision-benchHanno-Labs/decision-benchOpen benchmark runtime for document-grounded decision models.★ 1
    Python
  5. jev-synergy-screeningPistachioAIHQ/jev-synergy-screeningA Jev title-and-abstract screening experiment compared with Cohen Abstract Triage labels for an ADHD review.★ 1
    Python
  6. Jev-MedQAZtrura/Jev-MedQAExperiments running Jev on medical multiple-choice question answering.★ 1
    Python

See all 107 Benchmarks & research projects →

Where we found it