Gemini 3.5 Flash Lite (medium effort)
On the Tripstitch travel planning benchmark
Travel Score
86.4
Hallucination rate
20.3%
Route excess
8.6%
Constraints satisfied
100.0%
Gemini 3.5 Flash Lite (medium effort) is the first model on this benchmark, measured over 5 runs of 790 scored cases. Comparisons appear once a second model is tested.
- Model id
google/gemini-3.5-flash-lite- Reasoning effort
- medium
- Tested
- 22 July 2026
- Suite version
- 1.0
- Repeats
- 5
- Cases scored
- 790
- Cost per case
- $0.0005
- Median call
- 1.6s
- Usable output
- 100.0%
Against the field
Family scores for Gemini 3.5 Flash Lite (medium effort).
What a search tool is worth
Hallucination rate on identical questions, with and without a place-search tool.
Every measurement
Held out is the private half of the task set, never published. A large gap between the two columns is a sign the public cases leaked into training data.
| Measurement | Value | Spread | Held out | Rank |
|---|---|---|---|---|
| Place grounding 96 30% of score | ||||
| Hallucination rate place_existence Share of recommended places that do not exist: neither the places database nor Apple Maps can find anything by that name. | 20.3% | ± 5.5% | 16.0% | not scored |
| Wrong location rate place_existence Share of recommended places that are real but not in the city they were recommended for. | 7.0% | ± 1.2% | 5.3% | not scored |
| Requests fulfilled place_existence Share of the requested recommendations actually returned. Scored so that answering with fewer places cannot buy a better hallucination rate. | 100.0% | ± 0.0% | 100.0% | not scored |
| Hallucination rate grounded_place_existence Share of recommended places that do not exist: neither the places database nor Apple Maps can find anything by that name. | 0.3% | ± 0.7% | 0.0% | not scored |
| Wrong location rate grounded_place_existence Share of recommended places that are real but not in the city they were recommended for. | 0.0% | ± 0.0% | 0.0% | not scored |
| Requests fulfilled grounded_place_existence Share of the requested recommendations actually returned. Scored so that answering with fewer places cannot buy a better hallucination rate. | 100.0% | ± 0.0% | 100.0% | not scored |
| Fabrication rate abstention Share of impossible requests answered with invented places instead of an admission that none exist. | 0.0% | ± 0.0% | 0.0% | not scored |
| Honest refusals abstention Share of impossible requests where the model explicitly said no such places exist. | 100.0% | ± 0.0% | 100.0% | not scored |
| Geographic knowledge 86 20% of score | ||||
| Coordinate error coordinate_recall Median distance between the coordinates a model gives for a named place and where that place actually is. | 95 m | ± 22 m | 117 m | not scored |
| Coordinates within 250 m coordinate_recall Share of places a model locates closely enough to drop a usable map pin. | 63.6% | ± 0.0% | 75.0% | not scored |
| Distance error distance_estimate Median percentage error on the straight-line distance between two places. | 1.5% | ± 1.0% | 8.3% | not scored |
| Distance bias distance_estimate Mean signed error. Negative means the model consistently underestimates how far apart things are. | +2.9% | ± 1.9% | 7.7% | not scored |
| Spatial reasoning 67 15% of score | ||||
| Route excess over optimal route_optimization How much longer the model's visiting order takes to walk than the shortest possible order through the same stops. | 8.6% | ± 3.3% | 17.8% | not scored |
| Exactly optimal routes route_optimization Share of problems where the model found the shortest order outright. | 4.4% | ± 5.4% | 0.0% | not scored |
| Valid orderings route_optimization Share of answers that were a genuine permutation of the stops, with nothing dropped, duplicated or invented. | 100.0% | ± 0.0% | 100.0% | not scored |
| Spatial query accuracy spatial_queries Share of questions about supplied coordinates answered exactly right. | 71.1% | ± 5.4% | 80.0% | not scored |
| Real-world estimation 73 15% of score | ||||
| Travel time error travel_time_estimate Median percentage error against real road-network routing between two points. | 30.1% | ± 3.6% | 20.3% | not scored |
| Travel times underestimated travel_time_estimate Share of estimates that were optimistic by more than a tenth. Reported for direction, not scored: a model that always overestimates is not thereby good. | 65.0% | ± 3.3% | 100.0% | not scored |
| Arrival times exactly right itinerary_time_math Share of journeys where the local arrival time was correct to the minute. | 54.0% | ± 8.0% | 10.0% | not scored |
| Arrival time error itinerary_time_math Median distance between the local arrival time given and the real one. | 9.0 min | ± 14.5 min | 27.0 min | not scored |
| Itinerary construction 97 20% of score | ||||
| Constraints satisfied constraint_satisfaction Share of individually checkable constraints the produced plan actually satisfies. | 100.0% | ± 0.0% | 98.8% | not scored |
| Plans satisfying every constraint constraint_satisfaction Share of briefs where every single constraint held. One miss fails the plan, which is how a traveller experiences it. | 100.0% | ± 0.0% | 90.0% | not scored |
| Impossible legs schedule_feasibility Share of moves between consecutive stops where the gap in the schedule is shorter than the real walking time. | 17.7% | ± 3.9% | 25.0% | not scored |
| Time debt per day schedule_feasibility How many minutes short the day runs once every move is given the time it really takes. | 1.7 min | ± 0.3 min | 9.6 min | not scored |
| Budgets blown budget_arithmetic Share of plans whose own line items add up to more than the budget the brief set. | 0.0% | ± 0.0% | 0.0% | not scored |
| Totals that do not add up budget_arithmetic Share of plans where the total the model stated disagrees with the sum of the costs it listed. | 0.0% | ± 0.0% | 0.0% | not scored |
| Average overshoot budget_arithmetic How far over budget the plans that went over ended up, on average. | n/a | n/a | not scored | |
Full definitions of every measurement are on the methodology page. Back to the leaderboard.