Delivery app data and reviews: the playbook that separates before from after

Working delivery app data and reviews with a method lifts a rating from 4.2 to 4.7 in roughly eight weeks and lifts organic visibility inside the app with it, because the Rappi, Uber Eats, DiDi Food and iFood algorithms weight rating, prep time and cancellation rate above almost any other signal. The starting point is not replying to reviews: it is downloading the metrics panel, cross-referencing each complaint against the dish and the time slot that produced it, then fixing the cause in the kitchen. Replying without fixing moves perception for a week and sinks it again the following month. The sequence that works is five steps with a measurable deliverable each: baseline, cause map, kitchen fix, systematic public reply, weekly ranking control.
A quick-service operator in Medellín was billing 41 million pesos a month in Rappi delivery orders and lost 12 million of it in two months without touching price, menu or coverage area. The rating had slid from 4.6 to 4.1 after a run of orders arriving cold during the 8 p.m. peak, and the algorithm, which reads no excuses, pushed the listing from screen one to screen three. That is the whole mechanism: a review is not reputation, it is DISTRIBUTION.
Owners tend to treat the app panel as a report card rather than what it actually is, a ranking system with half-published rules, so they react to the aggregate rating when the lever sits three levels below it, in one specific dish arriving badly at one specific hour from one specific kitchen. Working delivery app data and reviews means dropping from that useless average down to actionable detail.
The gap between a dark kitchen and a brick and mortar restaurant shows up right here. A hidden kitchen has no storefront, no server to rescue a mediocre experience with a smile, no neighborhood recommending it by word of mouth: its only trust asset is the listing inside the app, with its stars, its photos and its comments. A ghost kitchen sitting at 4.3 stars is effectively closed half the time even with the lights on.
Side-by-side comparison
| BEFORE (no method) | AFTER (Masterestaurant method) | |
|---|---|---|
| Average in-app rating | ✕4.1 stars after 6 months trading | ✓4.7 stars by week 8 |
| Median position in category | ✕Rank 19 of 40 in the zone search | ✓Rank 4 of 40 in the same search |
| Share of reviews answered | ✕3% of comments, no criteria | ✓100% of negatives inside 24 hours |
| Orders cancelled by the venue | ✕6.4% of weekly volume | ✓1.1% of weekly volume |
| Declared vs real prep time | ✕Declares 18 minutes, delivers in 31 | ✓Declares 24 minutes, delivers in 22 |
| Average in-app ticket | ✕38,400 COP | ✓46,900 COP |
| Monthly delivery channel sales | ✕29 million COP | ✓48 million COP |
Step 1: pull the last 40 orders and forget the lifetime average
Download your last 30 days of order data and isolate the most recent 40 rated deliveries, because that short window is what moves your ranking inside the app while the cumulative average exists mainly to comfort the owner. A restaurant sitting on 1,200 reviews and a 4.6 lifetime score can be averaging 3.9 across its last 40 deliveries without the headline number budging a single decimal for six or seven weeks, and by the time the visible rating slides to 4.4 you have already lost two tiers of exposure. What this step delivers is a spreadsheet with four columns per order: exact time, dish, rating, and the customer's actual words. Nothing else. Verify it by counting rows: if you have fewer than 40 rated orders, widen the range to 60 days before drawing any conclusion, since below that floor noise beats signal every time. Sort those 40 rows by clock time and average each sixty-minute block, which is where the real pattern surfaces.
Step 2: cross rating against dispatch time to find the hour that is bleeding you
At the Medellín fast-food operation billing 41 million pesos a month through Rappi, the slide from 4.6 to 4.1 was not evenly spread: between 19:40 and 21:10 the average sat at 3.4 while the rest of the day held at 4.7, so the problem was never the kitchen but the kitchen AT PEAK. Twelve million pesos gone in two months over ninety badly handled minutes. Your deliverable here is a single bar per time slot, drawn by hand if that is faster. You will know it worked when you can point a finger at the guilty window and state how many orders it represents against the total, which typically lands somewhere near 30%. Any one-star review naming a dish, an hour, or a temperature is worth more than twenty silent five-star ratings, because it hands you free the audit a consultant would bill for.
Step 3: read the bad reviews hunting for the data point, not the insult
When a customer writes that the burger arrived cold at 21:40, that person is describing with surgical precision an assembly-time or packaging failure on the hot line during peak. The mistake I see over and over is answering that comment defensively and filing the person away as difficult. Tag every negative comment into one of four buckets: temperature, missing items, delay, food quality. Then tabulate. The deliverable is a plain count, and across most operations I review two buckets hold more than 70% of the complaints, which shrinks the whole problem down to two attackable causes instead of a vague feeling that service has gone soft. When temperature dominates, the fix almost never lives in the food and almost always lives in assembly order and packaging. Move the fries into a vented container, hold the bag build until the courier is under three minutes away, and raise your auto-accept threshold so the line never carries six simultaneous tickets it cannot sustain.
Step 4: attack the dominant cause with a measurable operational change, never a discount
Diego F. Parra keeps hammering at Masterestaurant that the prep time you declare in the app is a contract with the algorithm: declare 18 minutes and dispatch at 27, and the platform punishes your on-time rate long before any customer punishes your stars. Raise the declared time to whatever you actually hit. Losing three positions for promising speed costs more than showing up as an honest 25-minute kitchen, and 76% of US operators already treat technology as a competitive edge (National Restaurant Association, 2024) precisely because these adjustments can be measured. Ghost kitchens play by harsher rules, since they have no storefront, no server rescuing a mediocre meal with a smile, and no neighborhood passing the name around. Their only trust asset is the app listing: stars, photos, comments. A ghost kitchen at 4.3 is effectively closed half the day even with the lights on and the full crew clocked in.
Step 5: why a dark kitchen cannot afford a single tenth of a star
Flip it around for a second: if Rappi hid ratings tomorrow, the brick-and-mortar restaurant would keep selling through its window and word of mouth, while the dark kitchen would be reduced to fighting on price against something like 74,000 listings, roughly the universe DiDi Food runs in Mexico, where 70% are local small businesses (DiDi Food, 2024). That asymmetry explains why inside a ghost kitchen every tenth of a star gets defended like margin. The first and costliest is buying reviews or begging relatives for them: fraud systems at Uber Eats, which pushed close to 74.6 billion dollars in gross bookings during 2024 (Statista, 2024), spot clusters of fresh accounts with no order history and void the entire batch, leaving you worse off than when you began. Second, staring only at the aggregate score. Third, answering every comment with the same copied template, which any customer smells in two seconds.
Five mistakes that wreck this work before it starts
Fourth, switching the app off during a crisis believing that protects the score, when inactivity hurts placement as much as a bad star or more. Fifth, changing five things at once: move packaging, declared time, menu, and price in the same week and you will never know which one worked. One change, two weeks of measurement, then decide on data. Mark your start date and review the board every fourteen days with four numbers on the table. Recent-window rating should climb at least two tenths per fortnight until it settles near 4.7. Complaints in your dominant bucket should fall below 20% of all negative comments. Actual dispatch time in the critical slot should converge on the declared time with under four minutes of drift. And your own cancellation rate should sit below 2%. If three of those four move the right way by week four, you are on track and the job becomes maintenance.
How to know it all landed: the eight-week close?
If none of them moves, your time-slot diagnosis was wrong and you go back to step two with 60 orders instead of 40.
Global online food delivery is worth 1.51 trillion dollars in 2026 and grows at 6.24% a year (Statista, 2026); inside that tide, your spot on the first screen is the only thing actually up for negotiation. Delivery app data and reviews are worked on the LAST 30 to 100 orders, never on the full history. Rappi and Uber Eats weight the recent window because they want to know how you cook today, not how you cooked in March. A venue with 1,200 lifetime reviews and a 4.6 average can be falling off a cliff while its last 40 orders average 3.9, and the big number hides it for weeks. A negative review carrying a concrete detail beats ten loose stars, because it tells you exactly what to fix.
Four differences that decide the outcome
When a customer writes that the burger arrived cold at 9:40 p.m., they are handing you a free audit of your hot line at peak. The mistake I keep running into is treating that comment as a personal attack instead of the cheapest operational input available. Replying does not raise the rating; fixing does. The public reply works on the future reader deciding between your listing and the competitor next to it, and there it does move conversion: a specific, non-defensive answer turns a two-star review into a signal of a serious operator. Leave the cause sitting in the kitchen, though, and the next batch of reviews repeats the same text. In a dark kitchen the effect hits harder than in a brick and mortar restaurant. A hidden kitchen lives off in-app ranking with no street traffic to cushion it, so half a rating point translates straight into lost orders; a venue with a dining room absorbs that hit through reservations and walk-ins while it repairs the digital channel.
Before against after, criterion by criterion
What happens when nobody reads the panelBefore
- The rating drops 0.5 stars and the owner finds out through revenue, three weeks late.
- Complaints get read one by one, never grouped, so the pattern never surfaces: 60% name the same dish.
- Replies use the template the app suggests, identical across all 40 of them.
- Declared prep time is optimistic and drives cancellations the algorithm punishes twice over.
- Nobody crosses review against time slot, so the night-shift problem dissolves into the monthly average.
What changes once you read the dataMasterestaurant
- Weekly panel download, with rating segmented by dish, hour and courier.
- Every complaint sorted into four causes: temperature, missing items, packaging, delay. Attack the most frequent one.
- Public reply inside 24 hours, specific, naming the dish and the fix.
- Declared prep time carries six minutes of slack over the real time measured in the kitchen.
- Numeric checkpoint every Monday: if the last-30-orders rating falls under 4.5, freeze every promotion until it recovers.
Side-by-side comparison
| BEFORE (no method) | AFTER (Masterestaurant method) | |
|---|---|---|
| Average in-app rating | ✕4.1 stars after 6 months trading | ✓4.7 stars by week 8 |
| Median position in category | ✕Rank 19 of 40 in the zone search | ✓Rank 4 of 40 in the same search |
| Share of reviews answered | ✕3% of comments, no criteria | ✓100% of negatives inside 24 hours |
| Orders cancelled by the venue | ✕6.4% of weekly volume | ✓1.1% of weekly volume |
| Declared vs real prep time | ✕Declares 18 minutes, delivers in 31 | ✓Declares 24 minutes, delivers in 22 |
| Average in-app ticket | ✕38,400 COP | ✓46,900 COP |
| Monthly delivery channel sales | ✕29 million COP | ✓48 million COP |
The numbers behind this playbook
“We came in at 4.1 stars and 29 million monthly in the app, convinced the problem was Rappi's commission. We pulled the panel for our last 80 orders and 61% of the complaints named two dishes: the breaded chicken and the fries. We swapped the fries packaging for a vented one, pulled the breaded chicken off the menu between 8 and 10 p.m. because the line could not hold it, and raised declared prep time from 18 to 24 minutes. By week eight we sat at 4.7 and billed 48 million. We never touched price and never paid a peso in ads.”
Five steps, each with a deliverable and a numeric checkpoint
Before touching anything you need three accesses: the business panel of every app you sell on (Rappi Partners, Uber Eats Manager, DiDi Food, iFood Portal do Parceiro), your Google Business Profile listing and a plain spreadsheet. DELIVERABLE: a table with six figures per app — last-30-orders rating, lifetime rating, declared prep time, real prep time measured in the kitchen across five days, venue cancellation rate, and average position in your category search inside your zone, checked at 1 p.m. and at 8 p.m. Typical mistake: using the lifetime rating as the baseline, which almost always lies upward. CHECKPOINT: if you cannot fill all six cells per app, stop; half of these improvement plans fail for starting without real measurement.
Pull the last 80 to 100 comments from each app and sort them one by one into four buckets: temperature, missing items, packaging, delay. Add two more columns, dish mentioned and time slot. DELIVERABLE: a cause ranking ordered by frequency, with the highest-incidence dish and hour identified by name. In practice 55% to 70% of complaints cluster into two dishes and one slot. Typical mistake: sorting by gut feel instead of counting; do it with the sheet open and the number visible. CHECKPOINT: cause number one must explain at least 30% of negative complaints. If everything comes out evenly spread, widen the sample to 150 comments before concluding.
This is where it is won or lost, which is why it takes longer than the other four steps combined. Attack cause number one only, with one verifiable physical intervention: new packaging, moving a station on the line, temporarily pulling the problem dish during the critical slot, or raising declared prep time with six minutes of slack over the real figure. One intervention at a time, because if you change three things you will never know which one worked. DELIVERABLE: the intervention in place, with photo and date, plus real prep time measured again across five days. Typical mistake: gathering the team and asking for more care, which is a wish and not an intervention. CHECKPOINT: real time must land under declared time on 90% of the week's orders.
Answer 100% of reviews at three stars or below inside 24 hours, and one in three of the positives. Each reply carries three elements and none of them is a generic apology: the dish named, what got fixed specifically, and a concrete invitation to come back. Never argue, never repeat a phrase twice, never ask in writing for a rating change. DELIVERABLE: a 100% reply rate on negatives with average response time under 24 hours, verifiable in the panel. Typical mistake: the template the app suggests, identical across forty replies, which readers spot in two seconds and which subtracts credibility instead of adding it. CHECKPOINT: none of your last twenty replies shares more than six consecutive words with another.
Every Monday morning repeat the six measurements from step one, in under fifteen minutes, then search your category inside the app from a phone located in your delivery zone. DELIVERABLE: the week's row in the sheet, with search position logged at 1 p.m. and 8 p.m. One hard rule that will save you money: if the last-30-orders rating drops below 4.5, freeze every promotion and every geotargeted ad until it recovers, because paying to send traffic to a poorly rated listing accelerates the fall instead of braking it. CHECKPOINT: four consecutive weeks above 4.6 with position improving; if it stalls, go back to step two and widen the sample.
And with AI?
Optimize channels, pricing and unit economics of your dark kitchen. Diego F. Parra is an expert in AI applied to restaurants.
Free tools to apply this now
What to run alongside the execution
The steps above cost nothing except the owner's time, which is rarely free in practice. These three Masterestaurant tools keep the kitchen fix from eating the margin and keep the delivery channel profitable after commission, the part almost nobody calculates before jumping into selling on Rappi.
Order matters here: the cash number first, then the channel business model, and only at the end the scale. Diego F. Parra insists on reviewing food cost of your best-selling in-app dish before spending a peso on geotargeted ads, because a dish at 38% food cost with 27% commission does not get fixed by more orders, it gets worse.
Questions that land every week
How many reviews do I need to rank first on Uber Eats or Rappi?
How many reviews do I need to rank first on Uber Eats or Rappi?
There is no magic number and anyone promising you one is selling smoke. What the algorithm weights is the recent-window rating, compliance with declared prep time and cancellation rate. A venue with 90 reviews at 4.8 stars usually outranks one with 800 reviews at 4.2.
Can I ask a customer to delete a negative review?
Can I ask a customer to delete a negative review?
You can ask, and it is a bad idea. Apps penalize rating manipulation requests and customers routinely post a screenshot of the message, which multiplies the damage. Reply in public, fix the cause in the kitchen and let the next twenty good reviews dilute that one. It works faster than you would expect.
Why did my delivery app sales drop if I changed nothing?
Why did my delivery app sales drop if I changed nothing?
Something changed that you are not watching: real prep time stretched, a new competitor entered your zone, or the last-30-orders rating fell under 4.5 and the algorithm moved you off screen one. Measure those three before blaming commission or the season.
Is it worth answering five-star reviews too?
Is it worth answering five-star reviews too?
Yes, but one in three, and never with the same sentence. Answering a positive review speaks to the reader comparing listings, who sees somebody attentive behind the counter. Answering 100% of them with a template does the opposite: it reads like a bot and drains weight from the replies that matter, the negative ones.
Sector data 2026 (official sources)
Verifiable industry benchmarks from official, non-commercial sources (government, industry associations, market research) - not competitors.
| Metric | Benchmark 2026 | Source |
|---|---|---|
| IA para tomar pedidos de clientes en restaurantes de EE. UU. | Solo 6% de los restaurantes usa IA para tomar pedidos de clientes (2026) | National Restaurant Association 2026 |
| Restaurantes que ven la tecnología como ventaja competitiva EE. UU. | 76% de los operadores cree que la tecnología les da una ventaja competitiva | National Restaurant Association 2024 |
| Mercado global de cocinas en la nube en 2025 | USD 80.300 millones | Grand View Research — Cloud Kitchen Market 2025 |
| Proyección del mercado de cocinas en la nube a 2033 | USD 203.720 millones | Grand View Research — Cloud Kitchen Market 2033 |
| CAGR del mercado de cocinas en la nube 2026-2033 | 12,6% | Grand View Research — Cloud Kitchen Market |
| Cuota de Asia-Pacífico en cocinas en la nube 2025 | 48,0% de los ingresos | Grand View Research — Cloud Kitchen Market 2025 |
Related content
Grow your restaurant with the Masterestaurant method
Applied in +8.400 restaurants across 43 countries.
