jev-automation-testing
Mobile UI testing pipeline combining deterministic rules, Jev judgments and optional vision descriptions.
- GitHub stars
- 0 as of 2026-10-05
- Forks
- 0
- Language
- Python
- License
- MIT
- Created
- 2026-10-03
- Last updated on GitHub
- 2026-10-05
- Added to AionEdge
- 2026-10-05

What Jev does in jev-automation-testing
Jev judges text, extracted UI information and vision-model descriptions, with low-confidence results marked for review. Offline fake-driver fixtures are separate from live Appium and model runs. Catalog verification inspected README and source; the application was not independently runtime-tested.
Verification
Verified. We confirmed this repository through the GitHub API and located the code that calls Jev at a pinned commit: view the Jev code ↗. Upstream code changes over time, so check that the file still exists at the current commit before citing it.
Topics
- ai-agents
- mobile-testing
- screenshot-testing
- test-automation
How Evaluation & Observability projects use Jev
Tools for checking agent work and watching systems in production. They use Jev as a fast judge of whether conditions hold, so checks can run on every step instead of on a sample.
Related Jev projects in Evaluation & Observability
- foreman-jevShifty-Eye-Games/foreman-jevAn experimental Jev supervisor for Codex workers with programmer-selected acceptance commands.★ 0
Python
jev-eval4esv/jev-evalCompares Jev and OpenRouter models on labeled tasks for accuracy, calibration, latency and cost.★ 1
Python- jev-flash-reviewTheBous/jev-flash-reviewAn MCP review engine that evaluates Agent-supplied diffs against explicit rules.★ 1
TypeScript - hermes-jev-north-starpoponline63/hermes-jev-north-starA Hermes goal-checking skill that saves requirements, creates a run prompt and checks completion evidence.★ 1
Python - jev-playgroundmizchi/jev-playgroundA MoonBit and TypeScript Jev playground covering games, browsers, command risk and small languages.★ 19
TypeScript - goodwatch-monorepoalp82/goodwatch-monorepoA film-and-TV attribute-scoring experiment inside GoodWatch comparing Jev question designs and batch sizes.★ 38
Python