tripsnek
Optimize your Itinerary

TripSnek Lab

Decision Model vs. Conventional LLM - Building a natural language bridge to a formal travel planning system

We give Jev, a decision model that only ever answers multiple-choice questions, the same job as Claude: given real user descriptions of travel ideas, pull out what each traveler is asking for. With a few tricks it held its own, for of the cost. These are our notes, including the parts that were harder than we expected.

Daniel Tuohy · TripSnek ·

The short version

Jev matched Claude for of the cost

On real trip requests, Jev captures of what was asked for, at per 1,000 requests and each. Claude Sonnet captures , at and .

Jev pipelineClaude (LLM)

Accuracy against cost

Per 1,000 requests, log scale. Up and left is better.

Accuracy against time

Median seconds per request, log scale. Up and left is better.

Show as a table

But it needed as much code to make it work

What each pipeline needs on top of the lines of code they share. Each row has its own scale; lines of code leave out comments and blank lines.

ClaudeJev
Show as a table

Why we tried this

TripSnek plans multi-city trips around Europe. The core of it is an optimizer: a genetic algorithm that searches thousands of candidate itineraries (which cities, in what order, how many nights each, train or car or plane between them) and evolves them toward the best fit for what you asked for.

What you ask for comes in two kinds:

Constraints: must hold

  • Start and end city
  • Stops you want, and for how long
  • Dates, or total trip length
  • Places you've ruled out
  • Car, train, or a mix

Preferences: shape the score

  • Interests: food, wine, hiking, castles, beaches…
  • Pace: relaxed, moderate, or see-it-all
  • Flights between stops, or not

People don't think in constraint forms, though. They write "two weeks in October, landing in London and flying home from Rome, three nights in Paris, and we'd rather take trains." So we've been playing with a side project: a plain-language front door that reads a description like that and fills in the planner for you. It's an experiment, not a core part of TripSnek, and this page is about one branch of it: what happens if the model reading the post is Jev instead of an LLM?

Left, what you write: the describe-your-trip page with a paragraph asking for 18 days in Italy in May, into Rome and home from Venice, a few nights in Florence, maybe somewhere near the ocean, a car only if it's worth it, some museums, and plenty of food and wine. Right, what TripSnek plans from it: an 18-night itinerary of Rome 3 nights, Orvieto 1, Assisi 1, Montepulciano 2, Siena 2, Pisa on the way, Vernazza 3, Florence 3 and Venice 3, with a car picked up near Orvieto and dropped off near Siena, drawn as a route on a map of central and northern Italy.
The plain-language front door: a paragraph in, a planned trip out. Rome and Venice are the ends, and Florence and the coast at Vernazza get three nights each, as asked. The optimizer fills in the rest: a few days through Umbria and Tuscany (Orvieto, Assisi, Montepulciano, Siena), with a car for just that stretch.

The task: slot filling

If you want the textbook name, this is slot filling: read free text and fill a fixed set of typed fields. It's the same shape as a booking bot pulling "destination" and "date" out of "I need a flight to Boston on Friday." Our form just has more slots, and some of them repeat, like the list of stops.

Our test set is real posts from r/Europetravel, each hand-annotated with what it actually asks for. Here's a short one:

What they wrote

"Prague, Florence, Venice, Milan 2024. How would you schedule a full 14 days travel between these cities? Im also thinking of Pisa and Santorini."

What it asks for

stop Prague · Florence · Venice · Milan
stop Pisa (optional)
stop Santorini (optional)
totalNights 14
allowFlights true (Santorini)
interest MajorCities (low confidence)

Most posts are longer and messier, with typos, hedging ("maybe 2–2.5 weeks"), cities they've already seen, and places that aren't places ("Oktoberfest"). We score one thing: how much of what was asked for gets captured, feature by feature, weighted by how clearly the post states each one. Invented constraints cost points as well as missed ones. Turning that into a plannable trip is a separate step, and it's the same code whichever model did the reading.

Two very different tools for the job

Jev is a decision model. You send it a piece of text plus a list of typed questions, and it answers every question in one pass with a calibrated probability: a Choice over options you define, a Score on a rubric, or a Noul, the probability that a claim is true. What it can't do is write. It can't say "Lisbon"; it can only pick Lisbon from a list you gave it. Here's how that plays out on this problem:

LLM (Claude)Decision model (Jev)
What you sendThe post, a prompt and a JSON schemaThe post and a list of typed questions, each with its own options (about 40–60 per post)
What comes backA JSON document it wroteA probability for every option of every question
Places, dates, lengthsWrites them down however they appearCan only pick from spans your own code found in the text first
Yes/no and fixed choicesPicks oneA natural fit: one calibrated probability per option
Arithmetic, reading across sentencesComes freeNot available; do it in code, or go without
Made-up valuesPossible, since it can write anythingImpossible: every value is a span from the post or a label you defined
Something your code didn't anticipateThe model can still notice itInvisible
Where the judgment calls liveInside the modelIn your thresholds, one per question
Explains itselfYes, in a notes field, in plain wordsNo; you work backwards from the numbers
Cost per 1,000 posts (Sonnet)
Median time per post (Sonnet)
Code it needs of its own lines: a prompt, a schema and one API call lines of scanning, question-building and thresholds

Same job, same words

To keep the comparison fair, both pipelines are given the task in literally the same words. Every decision criterion lives in one file, and the same strings are interpolated into Claude's system prompt and attached to Jev's questions as the descriptions of their options. Change one and both change. Here's the definition of a ruled-out place, and how each model receives it:

The criterion

forbidden: "Somewhere they implicitly or explicitly rule out: they say they
  do not want it, indicate that one or more travelers have already seen or
  visited, or are explicitly skipping it."

In Claude's prompt

## What each place is to the trip

Every place plays exactly one of
these roles:
- **start** -- The trip begins here…
- **forbidden** -- Somewhere they
  implicitly or explicitly rule out…

As a Jev question

choice({
  task: "What is this place to the trip?",
  mention: "Vienna",
  context: "…we've already done Vienna…"
}, { start, end, stop,
     forbidden, none })
See more of the shared criteria
stop
Anywhere they have put on the table for this trip. This includes places they are only considering, asking about, or weighing against each other; undecided is not the same as unwanted. If they name a city while planning this trip and have not ruled it out, it belongs here.
none
Not part of this trip at all: the city they live in or fly from, a place on a past or future trip, a place named only as a comparison, or not a place at all. Not for a place they are merely undecided about.
priority: optional
They raise it as a candidate rather than a decision: weighing it against somewhere else, asking whether it is worth the time, floating it as a maybe. Leaving it out is a legitimate answer to what they asked.
pace: Moderate
They ask for both in the same breath: plenty of sights but without rushing. This is the common case, however strongly either half is worded.
transport: Mixed
Open to renting a car for part of it, or happy with whatever works.
claim: expects to fly
They expect to fly between stops within the trip. Flying into the first city and home from the last does not count.
claim: round trip
The trip begins and ends in the same city: they fly home from the city they arrived in, or use one city as a base for the whole trip.
claim: somewhere you'd stay
This names somewhere a traveler would base a stay (a town, city, island, valley, lake, coastal stretch or mountain area) rather than a single building, museum, monument, fountain, restaurant, festival or event that sits inside one of those.
interest: KidFriendly
They are travelling with children of 16 or under, or ask for things that suit them. "Kids" with no age given counts, and so does any stated young age; "kids" who turn out to be adults does not. A family trip is not on its own evidence of children.

Claude also gets the list of places, dates and lengths our scanner found in the post, as hints it's free to ignore, and both outputs go through the same validation and scoring code. A few rules exist only in Claude's prompt, for mistakes Jev can't make by construction, like writing a date in the wrong format.

What comes back

Here's the Prague post again, and what each model returned for it:

Claude Sonnet: a document

{
  "startLocation": { "name": "Prague" },
  "endLocation": { "name": "Milan" },
  "stopLocations": [
    { "name": "Florence" },
    { "name": "Venice" },
    { "name": "Pisa",
      "priority": "optional" },
    { "name": "Santorini",
      "priority": "optional" }
  ],
  "totalNights": 14,
  "notes": "Order Prague-Florence-
    Venice-Milan taken as start/end
    per listing; Pisa and Santorini
    mentioned only as ideas…"
}

Jev: 46 probabilities

Prague     role     stop 0.94
Milan      role     stop 1.00
Santorini  priority optional 0.95
tripStart           unstated 0.99
tripEnd             unstated 1.00
totalNights         "14 days" 1.00
pace                unstated 0.49
                    BlitzTour 0.42
allowFlights        0.51
interest_Ocean      0.50
… 36 more

Jev's numbers become the same kind of record only after our own code applies a threshold to each one, breaks ties, and fills in fallbacks. Here that gives six stops, two of them optional, 14 nights, and no start or end city. Nothing else clears its bar.

Sonnet's output reads like a person wrote it, notes included. Jev's is a pile of numbers, and every rule for turning them into a trip is ours to write. On this post the difference happened to favour Jev: Sonnet decided the first and last cities in the list must be the start and end, which the post never says, while Jev put 0.99 on "not stated". Across the corpus, that kind of thing goes both ways.

The easy parts and the hard parts

Some of the form mapped onto Jev's question types with no effort at all. The rest needed tricks, because the answer isn't one of a list you can write down in advance.

FieldShapeWith Jev
Flights between stops, round trip, 16 interestsYes/noEasyOne Noul claim each
Pace, transport, start monthOne of a fixed setEasyOne Choice each
How firmly they want each stopOne of a fixed set, per placeEasyonce the places exist
Start, end, stops, ruled-out placesOpen: any place, any spellingHardScan the text for candidates, then ask what each one is
DatesOpen: "Aug 15", "the 22nd", "mid October"HardParse candidates with regexes, then ask what each one marks
Trip length, nights per stopNumbers, often impliedHardScan phrases, offer number labels, and hope nothing needs adding up

The hard rows are where an LLM's robustness is easy to take for granted. It doesn't care how a traveler phrases something, it can do arithmetic, and it can connect a sentence at the top of a post to one at the bottom. With Jev, every one of those has to be something your code anticipated.

Free with an LLM, out of reach for Jev15 vs 20 nights
"…my wife and I are doing 5 night in Lisbon, 5 in Rome, and 5 in Munich. … I'm thinking of taking a night from the Lisbon portion and stopping somewhere between Rome and Munich… Venice, Bologna, Innsbruck…"
Claude15 nights, every run. It added up the stays, moved a night out of Lisbon, and said so in its notes.
JevThe post never says "15", so there's no span to pick. With no trip length to read, our code sums every stay it did read, including "a night or two" for each city they were only weighing up, and lands on 20.

That's the general shape of Jev's hardest failure: if the scanner doesn't surface something, it's invisible to everything downstream, and no question design gets it back. A phrasing the regexes don't know, a number that only exists after some arithmetic, a date written as "the week after Easter": an LLM shrugs these off, and for Jev each one is either another rule in the scanner or a guaranteed miss.

What made the difference

We started with the most obvious Jev pipeline and added one idea at a time, keeping everything else fixed. Three ideas did nearly all the work.

How much of each request was captured, one idea at a time

Each Jev row adds one idea to the row above. Claude rows are one structured-output call each, with the same downstream code. Dot = mean of 3 runs, line = run-to-run range.

Jev pipelineClaude (LLM)
Accuracy is agreement with a human annotation of what each post asks for, feature by feature.
Show as a table
Start

The naive version

Look up every city name we know in the post, ask Jev what role each one plays (start, end, stop, ruled out, or not part of the trip), and take its top answer. Every other field becomes a fixed-option question: trip length picked from "3 days … 7 weeks", pace from three labels, each interest a yes/no.

That gets , which is already more than Claude Haiku 4.5 manages (), though Haiku ran without thinking because it doesn't accept the adaptive-thinking setting the other Claude models used. The failures show what's missing. Dates vanish, because no question could return one. Places outside our gazetteer are never asked about. And trip length is picked from a menu even when the post says "14 days", or says nothing at all.

Idea 1

Build the menu from the text

If the model can only choose, the options should be exactly what the traveler wrote. A deterministic scanner finds every candidate span: places we know, capitalized names we don't, dates in any format we could think of, and duration phrases like "2-week" or "15 to 18 days". Jev's job shrinks from extraction to tagging. As a bonus, a made-up value becomes impossible: every place in the output is a substring of the post.

From the traces62% → 93%
"…looking to spend much of our time in the Alps region. The idea is to start the trip in Munich (brother wants to see Hofbrauhaus, BMW Museum, stop at Neuschwanstein Castle)…"
Before"Alps" isn't in our gazetteer, so it's never asked about, and the thing they most want disappears. Meanwhile the trip-length menu answers "14 days" for a post that gives no length.
AfterThe scanner offers unknown names too, and asks each one more question: is this somewhere you'd stay? Alps 0.95 · Hofbrauhaus 0.03 · BMW 0.06. The Alps become a stop; the beer hall and the museum don't.
Idea 2

Ask where the answer lives

Asking about each mention separately is right for stops, because a trip can have any number of them. It's wrong for things there's exactly one of. Asked "what is this place to the trip?", Jev reliably says a trip's last city is a stop, which is true, and almost never says it's the end. So the one-per-trip facts get asked once, about the whole trip: where does it start, where does it end, does it come back to where it began?

From the traces72% → 97%
"We've found that the most budget-friendly option for flights is a round trip to and from Amsterdam."
BeforeAsked mention by mention, Amsterdam reads as the start (0.60) and never as the end (at most 0.02). A round trip names its city once for both ends, so the trip comes back with no end city.
AfterOne question about the whole trip, "it returns to where it started", scores 0.51–0.59 across three runs and puts Amsterdam at both ends.

Two smaller rules of the same kind. Don't ask what the text can't answer: "how many nights here?" is only asked where a number appears near the mention, because otherwise Jev will happily read one off the neighbouring city. Ask positive claims: "are flights OK?" as one yes/no read silence as "no", and 17 of 30 trips came back ground-only. Splitting it into "expects to fly between stops" and "rules out flying" let silence mean silence.

Idea 3

Thresholds are product decisions

Jev gives you probabilities, so you get to choose where each line is drawn, and "take the top answer" leaves that choice unmade. We set each threshold by what a wrong answer costs. The output is a map the traveler adjusts: an extra city pinned on it costs one click to remove, but a city silently ruled out of the search is invisible. So a stop needs only a clear plurality (0.35 against a uniform 0.20), ruling a place out needs 0.6, and an interest, which quietly reweights the whole trip, needs 0.7. Where a threshold could be measured instead of argued, we measured it.

From the traces71% → 88%
"…will be arriving [in Zante] on the 22nd. Before that, we will have arrived in Greece on the 15th of August in Athens."
BeforeTaking the top answer gives Athens: 7 nights (0.67), which is the whole trip handed to one city, plus a night in Zante nobody mentioned.
AfterA number read out of a post has to clear 0.75, the level where every correct read in the corpus sat and the wrong ones didn't. Both stays are dropped, and the dates carry the trip instead.

We also tried a fourth idea, re-asking the answers that sit close to a threshold. It didn't help.1

Where it landed

With those three ideas, Jev captures of what was asked for, against for Claude Sonnet and for Claude Opus, which barely improves on Sonnet. Jev does it for of Sonnet's cost ( of Opus's), in about per post against . That's the first pair of charts at the top.

It's also steadier. Across three runs over the same posts, Jev's score moved by . The Claude models moved by (Opus), (Sonnet) and (Haiku).

And one failure caught our eye because it can't happen to Jev. In of Sonnet extractions, the output wasn't valid JSON, even though the request asked for schema-constrained output. That post scored zero for the run. A retry would usually fix it, but it's a kind of failure a decision model doesn't have: Jev never writes anything, so there's nothing to parse.

What it cost to build

Jev is cheaper to run. It was not cheaper to build. Both pipelines sit on the same lines of shared code: place grounding, validation and trip assembly. On top of that, Claude needs a prompt, an output schema and one API call. Jev needs a scanner to find the candidates, code to turn each one into questions, and the thresholds and assembly that turn answers per post back into a trip. The second chart at the top puts numbers on it.

The Jev side is the size of Claude's, but most of the difference is in kind rather than amount. About a third of the Claude side is the prompt, which is English. The Jev side has decision points against , and hand-set thresholds, each of which we had to pick by reading traces and then defend with a measurement.

Things to keep in mind

When this might be worth trying

We don't know yet how far this generalizes, but on this problem the pattern seemed to fit because:

TripSnek plans multi-city trips around Europe with an optimizer that balances the places you want against the time you have. The plain-language front door is still an experiment, but the planner is there to try.
Open TripSnek

Notes

1. Re-asking the close calls didn't help. Jev almost always picks the same label on the same input, but the probability behind it drifts: one decision, on unchanged text, scored 0.59, 0.40, 0.52, 0.50, 0.44, 0.48, 0.48 and 0.58 across eight runs, crossing its threshold five times. So we tried re-asking any answer within 0.2 of its threshold twice more and averaging the three readings. Over three full runs of the corpus the average came out at ( without it), and the run-to-run range narrowed only from to , while cost per 1,000 posts rose from to and median time from to . Single decisions really do flip, but the flips go both ways and cancel out across a corpus. It's now off by default. ↩