Shipping it at iFood
A real on-device AI feature in production
Meet iFood
A Brazilian food delivery app
- •
Users order food and groceries through the app
- •200 M+ customers
- •120 M+ orders/month
- •
After checkout, they follow the order in real time
- •
That experience is what we call Waiting
The waiting screen
A high-attention moment after checkout.
A simple user problem
What if the order is not for the user?
- •
Someone else may be waiting for the order
- •
They need status, ETA, address or pickup details
- •
Today, that conversation usually happens on WhatsApp
Version 1
A floating AI button at the bottom of the waiting screen.
Why it didn't ship
The model wasn't competing with another model. The feature was competing with the product.
- •
The bottom area already had an important product surface
- •
The FAB introduced a new competing interaction
- •
More clicks here could mean fewer clicks somewhere else
- •
The safest decision was not to launch it
Version 02
One button with one job. A much quieter place in the product.
Generating on-device
The loading state is actual local inference not a request waiting for our backend.
Completing the job
Success isn't generating text. It's helping the user send it.
Small models need constraints
A technically valid answer can still be completely wrong for the product.
- •
Greeted the restaurant instead of the recipient
- •
Mixed delivery vocabulary with pickup orders
- •
Returned internal status terminology
- •
Invented details when the promo left room for interpretation
Make the model boring
We don't want creativity here. We want a faithful transformation of known facts.
Can we show the button?
The feature only exists when inference can run right now.
Coverage is a product metric
Before asking who clicked the feature, we need to know who could even see it.
- •
Background download state
- •
Waiting version eligibility
Measure before rollout
At iFood scale, even limited device coverage means thousands of opportunities every day.
- •
100K+ eligible orders/day
What happened in production?
Local inference worked. But it wasn't instant.
- •
8.71 s median generation time
What we learned
Most of the hard problems weren't model problems.
- •
Start with one user job
- •
Placement matters as much as capability
- •
Hide is a valid fallback
- •
Constrain small models aggressively
- •
Measure coverage, not only inference
From demo to product
On-device AI is a product decision, not an SDK demo.