- Citation precision
- 0.673
- High-risk authoritative
- 0.458
- Claim coverage
- 0.290
Track 01 · VetEvidenceBench 2.5
More clinical claims you can actually audit.
Voyage delivered far more risk-weighted, citation-backed content across the same 50 veterinary cases. The advantage survives the prespecified competitor-favourable missing-evidence bound.
- Citation precision
- 0.507
- High-risk authoritative
- 0.064
- Favourable coverage bound
- 0.122
- Citation precision
- 0.319
- High-risk authoritative
- 0.020
- Favourable coverage bound
- 0.064
Track 02 · PetMemoryBench 2.0
The right memory.
For the right pet.
A multi-pet, high-interference workload probes current state, changed facts, cross-pet isolation, and compound continuity.
Under this synthetic workload, Voyage preserved pet-specific current state more reliably than the tested ChatGPT product snapshot. The selected-pet routing design favours Voyage; this is a dated product-system comparison, not a claim about every model or memory setting.
Read the full scorecard →Run it yourself
Inspect every input.
Reproduce every score.
Frozen prompts, dated baselines, adjudications, schemas, and dependency-free scorers are public. Test your own product against the same contracts.
git clone https://github.com/Voyage-Pet-AI/voyage-ai-vet-evals
cd voyage-ai-vet-evals
python -m pip install -e .
vet-evals score --track vet-evidence \
benchmarks/vet-evidence/results/adjudications-v2.5.json
vet-evals score --track pet-memory \
benchmarks/pet-memory/results/adjudications-v2.0.json
Freeze first
Use the published cohort, rubric, and aggregation without tuning.
Predict direction
Write the expected score direction before any score-moving change.
Report the system
Name the product surface, tools, memory settings, and capture date.