Alibaba

Qwen3.8 Max (medium effort)

Ranked 9 of 13 on the Tripstitch AI Travel Index

Travel Score

75.0

#9 of 13 ± 1.3 across 5 runs

Grounding

70.1%

#13 of 13 field median 84.5%

Route excess

10.5%

#7 of 13 field median 10.5%

Constraints satisfied

88.2%

#12 of 13 field median 98.6%

Qwen3.8 Max scores 75.0 on the Travel Score, behind most of the field. It recommends a real place 70.1% of the time from memory, inventing places more often than most. Its strongest family is place grounding at 86 of 100; its weakest is spatial reasoning at 53, against a field median of 59. At $0.0044 per scored case it is expensive relative to the field, with the slowest calls of any model here.

Model id
qwen/qwen3.8-max
Reasoning effort
medium
Tested
6 August 2026
Suite version
1.6
Repeats
5
Cases scored
733
Cost per case
$0.0044
Median call
41.5s
Usable output
98.7%
Tokens per call
3,788

Against the field

Family scores for Qwen3.8 Max (medium effort) next to the median of every model tested.

Accuracy by how known the place is

How far Qwen3.8 Max (medium effort)'s coordinates land from the real place, for famous landmarks, regional spots and obscure ones.

Travel Score against cost

Qwen3.8 Max (medium effort) highlighted against the rest of the field.

Every measurement

MeasurementValueSpreadField median
Place grounding 86 30% of score
Geographic knowledge 83 20% of score
Spatial reasoning 53 15% of score
Real-world estimation 66 15% of score
Itinerary construction 74 20% of score

Full definitions of every measurement are on the methodology page. Back to the leaderboard, or see what has changed in the changelog.