Google

Gemini 3.5 Flash Lite (medium effort)

On the Tripstitch travel planning benchmark

Travel Score

86.4

± 0.8 across 5 runs

Hallucination rate

20.3%

mean of 5 runs

Route excess

8.6%

mean of 5 runs

Constraints satisfied

100.0%

mean of 5 runs

Gemini 3.5 Flash Lite (medium effort) is the first model on this benchmark, measured over 5 runs of 790 scored cases. Comparisons appear once a second model is tested.

Model id
google/gemini-3.5-flash-lite
Reasoning effort
medium
Tested
22 July 2026
Suite version
1.0
Repeats
5
Cases scored
790
Cost per case
$0.0005
Median call
1.6s
Usable output
100.0%

Against the field

Family scores for Gemini 3.5 Flash Lite (medium effort).

What a search tool is worth

Hallucination rate on identical questions, with and without a place-search tool.

Every measurement

Held out is the private half of the task set, never published. A large gap between the two columns is a sign the public cases leaked into training data.

MeasurementValueSpreadHeld outRank
Place grounding 96 30% of score
Hallucination rate place_existence Share of recommended places that do not exist: neither the places database nor Apple Maps can find anything by that name.20.3%± 5.5%16.0%not scored
Wrong location rate place_existence Share of recommended places that are real but not in the city they were recommended for.7.0%± 1.2%5.3%not scored
Requests fulfilled place_existence Share of the requested recommendations actually returned. Scored so that answering with fewer places cannot buy a better hallucination rate.100.0%± 0.0%100.0%not scored
Hallucination rate grounded_place_existence Share of recommended places that do not exist: neither the places database nor Apple Maps can find anything by that name.0.3%± 0.7%0.0%not scored
Wrong location rate grounded_place_existence Share of recommended places that are real but not in the city they were recommended for.0.0%± 0.0%0.0%not scored
Requests fulfilled grounded_place_existence Share of the requested recommendations actually returned. Scored so that answering with fewer places cannot buy a better hallucination rate.100.0%± 0.0%100.0%not scored
Fabrication rate abstention Share of impossible requests answered with invented places instead of an admission that none exist.0.0%± 0.0%0.0%not scored
Honest refusals abstention Share of impossible requests where the model explicitly said no such places exist.100.0%± 0.0%100.0%not scored
Geographic knowledge 86 20% of score
Coordinate error coordinate_recall Median distance between the coordinates a model gives for a named place and where that place actually is.95 m± 22 m117 mnot scored
Coordinates within 250 m coordinate_recall Share of places a model locates closely enough to drop a usable map pin.63.6%± 0.0%75.0%not scored
Distance error distance_estimate Median percentage error on the straight-line distance between two places.1.5%± 1.0%8.3%not scored
Distance bias distance_estimate Mean signed error. Negative means the model consistently underestimates how far apart things are.+2.9%± 1.9%7.7%not scored
Spatial reasoning 67 15% of score
Route excess over optimal route_optimization How much longer the model's visiting order takes to walk than the shortest possible order through the same stops.8.6%± 3.3%17.8%not scored
Exactly optimal routes route_optimization Share of problems where the model found the shortest order outright.4.4%± 5.4%0.0%not scored
Valid orderings route_optimization Share of answers that were a genuine permutation of the stops, with nothing dropped, duplicated or invented.100.0%± 0.0%100.0%not scored
Spatial query accuracy spatial_queries Share of questions about supplied coordinates answered exactly right.71.1%± 5.4%80.0%not scored
Real-world estimation 73 15% of score
Travel time error travel_time_estimate Median percentage error against real road-network routing between two points.30.1%± 3.6%20.3%not scored
Travel times underestimated travel_time_estimate Share of estimates that were optimistic by more than a tenth. Reported for direction, not scored: a model that always overestimates is not thereby good.65.0%± 3.3%100.0%not scored
Arrival times exactly right itinerary_time_math Share of journeys where the local arrival time was correct to the minute.54.0%± 8.0%10.0%not scored
Arrival time error itinerary_time_math Median distance between the local arrival time given and the real one.9.0 min± 14.5 min27.0 minnot scored
Itinerary construction 97 20% of score
Constraints satisfied constraint_satisfaction Share of individually checkable constraints the produced plan actually satisfies.100.0%± 0.0%98.8%not scored
Plans satisfying every constraint constraint_satisfaction Share of briefs where every single constraint held. One miss fails the plan, which is how a traveller experiences it.100.0%± 0.0%90.0%not scored
Impossible legs schedule_feasibility Share of moves between consecutive stops where the gap in the schedule is shorter than the real walking time.17.7%± 3.9%25.0%not scored
Time debt per day schedule_feasibility How many minutes short the day runs once every move is given the time it really takes.1.7 min± 0.3 min9.6 minnot scored
Budgets blown budget_arithmetic Share of plans whose own line items add up to more than the budget the brief set.0.0%± 0.0%0.0%not scored
Totals that do not add up budget_arithmetic Share of plans where the total the model stated disagrees with the sum of the costs it listed.0.0%± 0.0%0.0%not scored
Average overshoot budget_arithmetic How far over budget the plans that went over ended up, on average.n/an/anot scored

Full definitions of every measurement are on the methodology page. Back to the leaderboard.