OpenAI
GPT 5.6 Terra (medium effort)
Ranked 2 of 9 on the Tripstitch AI Travel Index
Travel Score
87.5
Grounding
88.8%
Route excess
3.6%
Constraints satisfied
100.0%
GPT 5.6 Terra scores 87.5 on the Travel Score, among the strongest tested. It recommends a real place 88.8% of the time from memory, inventing fewer places than most. At $0.0036 per scored case it is expensive relative to the field.
- Model id
openai/gpt-5.6-terra- Reasoning effort
- medium
- Tested
- 23 July 2026
- Suite version
- 1.5
- Repeats
- 5
- Cases scored
- 733
- Cost per case
- $0.0036
- Median call
- 10.3s
- Usable output
- 100.0%
- Tokens per call
- 2,829
Against the field
Family scores for GPT 5.6 Terra (medium effort) next to the median of every model tested.
Accuracy by how known the place is
How far GPT 5.6 Terra (medium effort)'s coordinates land from the real place, for famous landmarks, regional spots and obscure ones.
Travel Score against cost
GPT 5.6 Terra (medium effort) highlighted against the rest of the field.
Every measurement
| Measurement | Value | Spread | Field median |
|---|---|---|---|
| Place grounding 98 30% of score | |||
| Grounding, from memory Place existence higher is better Share of recommended places whose names resolve, meaning the places database or Apple Maps can find a place by that name. A name that resolves nowhere is counted as invented. Asked with no tool available, so the model is answering from memory. | 88.8% | ± 2.3% | 84.5% |
| Wrong location rate, from memory Place existence lower is better Share of recommended places whose names resolve, but only outside the city they were recommended for. Asked with no tool available, so the model is answering from memory. | 4.8% | ± 1.8% | 6.4% |
| Requests fulfilled, from memory Place existence higher is better Share of the requested recommendations actually returned. Scored so that answering with fewer places cannot buy a better grounding rate. Asked with no tool available, so the model is answering from memory. | 100.0% | ± 0.0% | 100.0% |
| Grounding, with search Place existence (with search) higher is better Share of recommended places whose names resolve, meaning the places database or Apple Maps can find a place by that name. A name that resolves nowhere is counted as invented. Measured on the identical questions with a place-search tool available, so the gap against the from-memory figure is what looking things up is worth. | 100.0% | ± 0.0% | 99.7% |
| Wrong location rate, with search Place existence (with search) lower is better Share of recommended places whose names resolve, but only outside the city they were recommended for. Measured on the identical questions with a place-search tool available, so the gap against the from-memory figure is what looking things up is worth. | 0.0% | ± 0.0% | 0.0% |
| Requests fulfilled, with search Place existence (with search) higher is better Share of the requested recommendations actually returned. Scored so that answering with fewer places cannot buy a better grounding rate. Measured on the identical questions with a place-search tool available, so the gap against the from-memory figure is what looking things up is worth. | 100.0% | ± 0.0% | 98.9% |
| Fabrication rate Abstention lower is better Share of impossible requests answered with place names that do not resolve in the area, instead of an admission that none exist. | 0.0% | ± 0.0% | 0.0% |
| Honest refusals Abstention higher is better Share of impossible requests where the model explicitly said no such places exist. | 100.0% | ± 0.0% | 100.0% |
| Geographic knowledge 88 20% of score | |||
| Coordinate error Coordinate recall lower is better Median distance between the coordinates a model gives for a named place and where that place actually is. | 49 m | ± 6 m | 49 m |
| Coordinates within 250 m Coordinate recall higher is better Share of places a model locates closely enough to drop a usable map pin. | 75.4% | ± 3.1% | 64.6% |
| Distance error Distance estimation lower is better Median percentage error on the straight-line distance between two places. | 0.1% | ± 0.0% | 0.5% |
| Distance bias Distance estimation closest to zero is best Mean signed error. Negative means the model consistently underestimates how far apart things are, which is the direction that schedules a day nobody can walk. | +0.9% | ± 1.8% | 0.9% |
| Spatial reasoning 76 15% of score | |||
| Route excess over optimal Route optimization lower is better How much longer the model's visiting order takes to walk than the shortest possible order through the same stops. An answer that is not a real ordering is recorded at the ceiling. | 3.6% | ± 0.6% | 4.3% |
| Route excess when the order is real Route optimization lower is better How much longer the route takes than the best possible, counting only answers that were a genuine ordering of the stops. Reported, not scored. | 3.6% | ± 0.6% | 4.3% |
| Exactly optimal routes Route optimization higher is better Share of problems where the model found the shortest order outright. | 12.7% | ± 7.3% | 9.1% |
| Valid orderings Route optimization higher is better Share of answers that were a genuine permutation of the stops, with nothing dropped, duplicated or invented. | 100.0% | ± 0.0% | 100.0% |
| Spatial query accuracy Spatial queries higher is better Share of questions about supplied coordinates answered exactly right. | 96.4% | ± 4.5% | 96.4% |
| Real-world estimation 76 15% of score | |||
| Travel time error Travel time estimation lower is better Median percentage error against real road-network routing between two points. | 25.2% | ± 4.1% | 23.0% |
| Travel times underestimated Travel time estimation lower is better Share of estimates that were optimistic by more than a tenth. Reported for direction, not scored: a model that always overestimates is not thereby good. | 58.7% | ± 5.0% | 65.3% |
| Arrival times exactly right Time zone arithmetic higher is better Share of journeys where the local arrival time was correct to the minute. | 90.0% | ± 3.3% | 90.0% |
| Arrival time error Time zone arithmetic lower is better Average distance between the local arrival time given and the real one, counting only the journeys it got wrong once the exact ones are averaged in. | 6.0 min | ± 2.0 min | 8.0 min |
| Itinerary construction 89 20% of score | |||
| Constraints satisfied Constraint satisfaction higher is better Share of individually checkable constraints the produced plan actually satisfies. | 100.0% | ± 0.0% | 98.9% |
| Plans satisfying every constraint Constraint satisfaction higher is better Share of briefs where every single constraint held. One miss fails the plan, which is how a traveller experiences it. | 100.0% | ± 0.0% | 92.7% |
| Legs that fit Schedule feasibility higher is better Share of moves between consecutive stops where the schedule leaves at least the real walking time. The rest are moves nobody could make. | 57.1% | ± 4.8% | 48.9% |
| Time debt per day Schedule feasibility lower is better How many minutes short the day runs once every move is given the time it really takes. | 7.2 min | ± 1.6 min | 9.4 min |
| Budgets blown Budget arithmetic lower is better Share of plans whose own line items add up to more than the budget the brief set. | 0.0% | ± 0.0% | 0.0% |
| Totals that do not add up Budget arithmetic lower is better Share of plans where the total the model stated disagrees with the sum of the costs it listed. | 0.0% | ± 0.0% | 0.0% |
| Average overshoot Budget arithmetic lower is better How far over budget the plans that went over ended up, on average. | n/a | 2.3% | |
Full definitions of every measurement are on the methodology page. Back to the leaderboard, or see what has changed in the changelog.