Global flight routes connecting destinations across a detailed navy globe
INDEPENDENT EVALUATION. REAL-WORLD TRAVEL.

The evaluation lab for AI travel agents.

AI assistants now plan and book trips for millions of travelers. We are grading them — with open benchmarks built from real traveler requests and real itineraries.

Built on travel expertise. Designed for real-world trust.
BETTER MEASUREMENT. BETTER TRAVEL.
Open benchmarks Expert-built tasks Independent evaluation
GRADING, NOT GUESSING

Agents are booking travel.
We are grading them.

Travel is more than a search result. It takes accuracy, good questions, and judgment — so we test for all three, and publish what we find.

Flights that don’t exist.

Agents invent flights and fares. A convincing answer is not the same as a bookable itinerary.

Assumptions instead of questions.

They skip the clarifying questions a good travel agent would ask — and guess what the traveler needs.

The cheapest isn’t always right.

They optimize for price, not the right fit: the cabin, the connection, or the needs of the traveler.

THE BENCHMARK FAMILY

One benchmark per travel vertical.

Different journeys. Different challenges. A dedicated test for each.

Starting with flights
Roadmap

awardflightsbench.

Points-and-miles redemptions: program rules, phantom availability, and fuel surcharges.

Roadmap

hotelbench.

Hotel search and booking: room types, policies, and rate integrity.

Roadmap

dmcbench.

Destination management: ground operators, transfers, tours, and local logistics.

Roadmap

villabench.

Luxury villas and high-end stays: the judgment-heavy segment.

THE METHODOLOGY

How do we grade agents?

Not just “did it answer?” — but “did it understand, deliver, and earn the traveler’s trust?”

01

Real client conversations

Tasks come from genuine, sanitized traveler requests — not invented examples.

02

Hidden-context grading

The grader knows the traveler’s needs; the agent must uncover them by asking, not guessing.

03

Two separate verdicts

Task completion gets a capability score. Trusted completion gets a pass or fail: no invented flights or fares.

04

A living benchmark

Tasks are refreshed quarterly, so agents can’t simply train to the test.

RESULTS, NOT CLAIMS

Leaderboard.

A clearer picture of what AI travel agents can actually do. Open results, grounded in real tasks.

Coming soon

First flightsbench results publishing soon.

No rankings until the evaluations are in.

Get in touch for updates
A FEW GOOD QUESTIONS

What should you know about TravelLabs?

Straight answers about what we’re building, how it works, and how to get involved.

Travel expertise. Applied to AI.

TravelLabs was founded by a career travel-industry professional with deep expertise in complex flight itineraries. We’re building the measurement layer AI travel needs — combining real-world travel judgment with rigorous, independent evaluation.

LET’S BUILD A HIGHER STANDARD

Better AI travel starts
with better evaluation.

hello@travellabs.ai

For AI labs & enterprises

Private benchmark licensing and managed evaluations.
Know where your agents stand before your travelers do.

Work with us

For travel experts

Join our paid task-authoring network.
Bring your expertise to the next generation of travel agents.

Contribute your expertise