aionedge.
Jev project · Evaluation & Observability

jev-automation-testing

vankhangfet/jev-automation-testing

Mobile UI testing pipeline combining deterministic rules, Jev judgments and optional vision descriptions.

View on GitHub ↗See the Jev code ↗
GitHub stars
0 as of 2026-10-05
Forks
0
Language
Python
License
MIT
Created
2026-10-03
Last updated on GitHub
2026-10-05
Added to AionEdge
2026-10-05
Built by this developer
vankhangfet

What Jev does in jev-automation-testing

Jev judges text, extracted UI information and vision-model descriptions, with low-confidence results marked for review. Offline fake-driver fixtures are separate from live Appium and model runs. Catalog verification inspected README and source; the application was not independently runtime-tested.

Verification

Verified. We confirmed this repository through the GitHub API and located the code that calls Jev at a pinned commit: view the Jev code ↗. Upstream code changes over time, so check that the file still exists at the current commit before citing it.

Topics

How Evaluation & Observability projects use Jev

Tools for checking agent work and watching systems in production. They use Jev as a fast judge of whether conditions hold, so checks can run on every step instead of on a sample.

Related Jev projects in Evaluation & Observability

  1. foreman-jevShifty-Eye-Games/foreman-jevAn experimental Jev supervisor for Codex workers with programmer-selected acceptance commands.★ 0
    Python
  2. jev-eval4esv/jev-evalCompares Jev and OpenRouter models on labeled tasks for accuracy, calibration, latency and cost.★ 1
    Python
  3. jev-flash-reviewTheBous/jev-flash-reviewAn MCP review engine that evaluates Agent-supplied diffs against explicit rules.★ 1
    TypeScript
  4. hermes-jev-north-starpoponline63/hermes-jev-north-starA Hermes goal-checking skill that saves requirements, creates a run prompt and checks completion evidence.★ 1
    Python
  5. jev-playgroundmizchi/jev-playgroundA MoonBit and TypeScript Jev playground covering games, browsers, command risk and small languages.★ 19
    TypeScript
  6. goodwatch-monorepoalp82/goodwatch-monorepoA film-and-TV attribute-scoring experiment inside GoodWatch comparing Jev question designs and batch sizes.★ 38
    Python

See all 9 Evaluation & Observability projects →

Where we found it