On this page
This company published a study in this exact area, so it is fair to hold us to the same standard. What follows applies to our work as much as to anyone else's.
Three different things
Behaviour: strong evidence, narrow scope
Whether a guest returned, how often, what they ordered, how spend changed over time. This is directly observed, requires no inference about mental states, and is already sitting in the point of sale.
Its limitation is that it tells you what happened and not why. A guest who stopped coming may have moved house.
Text: moderate evidence, broader scope
What guests write can be classified into themes reasonably reliably. Our study did this at scale: 13,063 reviews across 41 Texas venues, classified twice by separate models, then checked against state alcohol tax filings so the classification had an external reference rather than only internal consistency.
Two independent models agreeing is meaningfully better than one model asserting. It is still weaker than a controlled experiment, and reviewers are not a random sample of guests. People who write reviews are unusual in ways that matter, and any conclusion has to carry that caveat rather than mention it once in a footnote.
Feeling: not measured, frequently claimed
A sentiment score is a classification of text, presented as a measurement of an internal state. Those are different things. A guest who wrote a lukewarm review may have had a fine evening and be a lukewarm writer.
The tell is precision. A product reporting that guest sentiment is 72.4 out of 100 has taken a fuzzy classification and given it a decimal point, and the decimal point is a design decision rather than a finding.
The question that separates good products from bad
Ask what the product cannot see.
A vendor with a clear answer has thought about the boundary of their method. A vendor who cannot name a limitation either has not thought about it or would prefer not to discuss it, and both are informative.
For our own study, the answer is specific. It cannot tell you what any individual guest felt. It cannot see guests who left without writing anything, which is most of them. It describes relationships and does not establish causes, so a venue with better reviews and better trading has not been shown to have one because of the other.
What to actually use
The practical stack, in descending order of how much weight it will bear:
Behaviour from the transaction record. Return frequency, spend, mix, timing. Direct, already yours, and underused.
Operational proxies joined to the shift. Waits, comps, remakes, complaints, attached to a specific night with a specific staffing pattern. Indirect but diagnosable, which is what makes it actionable.
Text themes at volume. Useful for direction and for noticing a change. Not useful at the level of an individual review, where the sample is one.
Sentiment scores. Treat as a rough directional indicator and ignore the decimal.
Why this matters commercially
A venue that acts on an overstated measurement wastes effort on a problem that may not exist, and loses confidence in measurement generally when the effort produces nothing. That is the real cost of the overselling in this category: not the licence fee, but that it teaches operators that analytics does not work.
Our method is written out on the research page so it can be judged rather than taken on trust. CoreTAP works on the first two items in the list above, which are the ones that carry weight.
