You.com Search Evals
Automated Search Evals

Provider Benchmark Dashboard

You.com · Evals

Each search provider is benchmarked the way an AI agent uses it — its results feed a fixed answer model (GPT⁠-⁠5.6 Luna), and the answers are graded across five public QA benchmarks. Every You.com sampler is highlighted; lower time and cost, and higher accuracy, is better.

Benchmark

You.com Evals
Providers
Model time = answer-generation wall clock; search time = retrieval wall clock (p50 across questions). Rendered by the automated eval pipeline from the latest Braintrust runs.