On this page
Three findings that surprised us
Guests who name a staff member in a review are far more positive. That result replicated three separate times across our earlier volumes. It is also, for revenue purposes, close to irrelevant: naming an employee separates high-revenue venues from lower-revenue ones by a ratio of just 1.20x.
What does separate them is occasion language. Reviews that mention a reservation appear 3.40x more often at the highest-revenue venues. Celebration language runs 2.81x. Group and party-size language 2.25x. The guest who books is worth more than the guest who is delighted, because booking converts a discretionary visit into a scheduled obligation.
This is the finding we most expected to be wrong, and it survived every robustness check we ran. The overall rate at which guests complain about something is statistically indistinguishable from zero as a predictor of receipts.
That does not mean complaints are free. Broken out by aspect, food complaints do track revenue (r = -0.489 on the consensus label set). It is the aggregate negativity that carries no signal. Read the caveat below before acting on this.
Operators tend to look for the single thing to fix. In 32 of 41 venues, the top two complaint categories sat within 25 percent of each other on mention count. There was no single thing. That has direct consequences for how you prioritise, and it is the kind of result you only get by counting every review rather than asking someone to summarise a sample.
The caveat that matters most
Every venue in this study is a top-10 alcohol earner in its city. All 41 of them are winners already, so every correlation here is a comparison among winners and none of it is causal.
The finding that negatives carry no revenue cost applies inside that elite cohort. It says nothing about a venue with a 29 percent negative rate that is not already a top earner, because venues like that are absent from the sample precisely because they failed. If you take one caution from this page, take that one.
How it was measured
Three tiers, deliberately kept separate so that the exact parts stay exact.
- State tax filings. TABC Mixed Beverage Gross Receipts, dataset
naix-2893. Compulsory, audited, monthly, venue level, public. - Star-derived sentiment. Computed in Python with no model involvement. Identical across every model tested.
- Aspect labels. All 13,063 reviews classified individually by two frontier models. 98.1 percent polarity agreement across 23,993 aspect pairs. All counting, percentages and ranking done deterministically in code.
The full method, including what we got wrong and had to rebuild, is here.
The volumes
- What drives restaurant revenue, the commercial argument across all 41 venues.
- Exact measurement method, the authoritative volume on how every number was produced.
The wider program runs to six volumes and roughly 63,000 words, covering the hospitality literature, guest and worker language, and city-level breakdowns for Houston, Katy, Sugar Land and The Woodlands.
Data ethics
Review scraping ran with personal data collection disabled throughout. No reviewer names, identifiers, photographs or profile links were retrieved or stored. Reviews are quoted by star rating and date only. Employee first names that appear are ones guests published publicly, and they are reported only as aggregate patterns, never attached to an individual's performance.
One venue's review corpus was excluded after we determined it was not authentic. Its tax filings remain in the receipts analysis; its review data is not used anywhere.
