Putting The Genie
in the Bottle

A Crash Course on running

LLMs on Android

@iurysza
iurysouza.dev

About me

  • •
    GDE for Android
  • •
    Senior Engineer @ Klarna
Iury Souza on stage
QR code for iurysouza.dev
@iurysza
iurysouza.dev

About me

  • •
    SDE Android
  • •
    Order Experience @ iFood
Gabriel Volpi on stage
@_gvolpi
@iFood

But why tho?

Do LLMs on mobile make any sense?

But why tho?

  • •
    Aren't LLMs super hardware intensive?
  • •
    We already have ChatGPT and Gemini in the cloud
  • •
    SOTA models get all the attention
  • •
    Benchmarks get crushed every week
WHAT GETS BURIED IN THE NEWSThe model icebergGemini 3.8 FlashClaude Opus 5.5Claude Fable 5.1DeepSeek V4 ProDeepSeek V4.1 FlashSOTA models get the attentionGemma 4 E2BQwen3.5 4BGranite 4.2 3BGemma 4 12BSmolLM3Efficient open models, mostly unseen
TOP LEFT CORNER MODELS
Another category of breakthroughs is happening in parallel
Highly efficient, capable open models
Gemma 4 31B matches Gemini 2.5 Pro on GPQA Diamond

Why run locally?

Even if they're getting good, why should you still bother?

  • •
    Performance and latency
  • •
    Privacy and security
  • •
    Cost and accessibility
  • •
    Offline functionality
RUNNING ON THE EDGETwo ways to run an LLM on deviceOPTION 1A system-wide modelOne model, shared by every appGemini Nano, built into AndroidCalled via ML Kit GenAI or Prompt APIThe platform keeps it updatedOPTION 2A model per appEach app ships its ownYou pick the modelYou bundle and run itWider device compatibility

System-wide LLM

One model, shared by every app

One model, shared by every app

Gemini Nano on Android

ML Kit GenAI and the Prompt API

  • •
    One model, shared by all apps
  • •
    The platform manages model updates
  • •
    Built-in safety features
  • •
    Feature APIs are the easy way in. The Prompt API gives you more control
  • •
    Supported devices only, checked at runtime
GEMINI NANORequirements1A compatible deviceNano v4: Pixel 11 series, Galaxy Z Flip8 & Z Fold8Nano v3: Pixel 9 & 10 series, Galaxy S26 seriesOnePlus 15, OPPO Find X9, vivo X300, Xperia 1 VIIINano v2: Galaxy Z Fold7, OnePlus 13, Xiaomi 15 & 17About 85 phones and tablets in total2AICore APKThe Android AICore appHosts and updates the model3Private Compute Services APKThe Private Compute Services appDelivers model updatesSource: ML Kit GenAI · Prompt API device support · Oct 2026

Using it in your app

ML Kit Prompt API

implementation("com.google.mlkit:genai-prompt:1.0.0-beta4")
val generativeModel = Generation.getClient()
val status = generativeModel.checkStatus()
generativeModel.generateContentStream(prompt)
.collect { chunk -> print(chunk.candidates[0].text) }

With it you can

ML Kit Prompt API

  • •
    Check if the device is supported, or the model still needs to download
  • •
    Choose a model variant
  • •
    Stream the response
  • •
    Set temperature, topK and max output tokens
GEMINI NANO · HOW DOES IT WORK? ROUGHLYYour appML Kit GenAIAICoreLoRAGemini NanoSafety featuresTPU / NPUHardware acceleratorAICOREOS-level integrationModel deployment managementHardware abstraction layerCloud serviceModel downloadPrivate Compute ServicesPCS + PCC: privacy through data isolationPCC: processes data safelyPCS: model updatesOne model for every app

ML Kit GenAI APIs

The high-level route to Gemini Nano

  • •
    High-level APIs for:
    • •Summarization
    • •Proofreading
    • •Rewriting
    • •Image description

Adding it to your app

ML Kit GenAI APIs

val articleToSummarize = "We are excited to announce a set of on-device..."
val options = SummarizerOptions.builder(context)
.setInputType(InputType.ARTICLE)
.setOutputType(OutputType.ONE_BULLET)
.setLanguage(Language.ENGLISH)
.build()
val summarizer = Summarization.getClient(options)
RUNNING ON THE EDGETwo ways to run an LLM on deviceOPTION 1A system-wide modelOne model, shared by every appGemini Nano, built into AndroidCalled via ML Kit GenAI or Prompt APIThe platform keeps it updatedOPTION 2A model per appEach app ships its ownYou pick the modelYou bundle and run itWider device compatibility

Self managed models

The DIY alternative

The DIY alternative

  • •
    You're in control
  • •
    Model management, setup, initialization and resource usage are yours
  • •
    Run many different models
  • •
    Bring your own fine-tuned model
  • •
    Wider device compatibility

LiteRT-LM

Adding it to your app

What model do we use?

  • •
    Export a Hugging Face model to a single .litertlm file with litert-torch
  • •
    The file bundles the weights, the tokenizer and the prefill and decode graphs
  • •
    Wait, what? 🤨

What's under the hood

What are these models?

  • •
    Gemma 4, from Google
  • •
    Open weights, Apache 2.0
  • •
    E2B and E4B are the phone-sized ones
  • •
    What are these models?
WHAT'S UNDER THE HOODModel familyabout 2B effective paramsinstruction-tunedLiteRT-LM bundle❯adb pushgemma-4gemma-4E2BE2Bitit.litertlm.litertlm--Source: Hugging Face · litert-community/gemma-4-E2B-it-litert-lm · 2.58 GB
TOKENIZER_MODELText to token IDsThe vocabulary maps each word to an IDToken IDs back to textTokenizer123456789explain123this456model789vocabulary (simplified){"explain":123,"this":456,"model":789}

Prefill and decode

What's under the hood

  • •
    Model weights and architecture
  • •
    LiteRT
  • •
    Prefill and decode operations

What are they doing?

What's under the hood

  • •
    What do we do with that?
  • •
    We're running inference
  • •
    Making next token predictions
INFERENCE FLOWA prompt goes in. Everything on the left runs oncePrefill reads the whole prompt and fills the KV cacheDecode: one token per lap, the prompt is never re-readEOS? Yes. The loop ends and the text is returnedPREFILL, ONCEPromptgenerate() calledTokenizetext to token IDsPrefillwhole prompt at onceKV cacheinitial stateDECODE LOOP, ONE TOKEN PER LAPNoYesnext position, updated KV cacheDecode stepdecode signatureLogitstoken scoresGreedy sampletop score winsDetokenizeID to textAppendadd to outputEOS?OUTPUTexplain123this456model789Return textSource: google-ai-edge/mediapipe-samples, litert_inference/Gemma3_1b_fine_tune.ipynb

Where were we again?

What model do we use?

Adding it to your app

LiteRT-LM

implementation("com.google.ai.edge.litertlm:litertlm-android:0.17.1")

Push the model to the device

LiteRT-LM

# Push the model to the device
MODEL_FILE="gemma-4-E2B-it.litertlm"
TARGET_DIR="/storage/emulated/0/Android/data/com.myapp/files/"
adb push "$MODEL_FILE" "$TARGET_DIR"

Create the engine

LiteRT-LM

val modelPath = File(
context.getExternalFilesDir(null),
modelName
).absolutePath
val engine = Engine(
EngineConfig(modelPath = modelPath, backend = Backend.GPU())
)
engine.initialize() // slow: call it off the main thread

Start a conversation

LiteRT-LM

val conversation = engine.createConversation(
ConversationConfig(
samplerConfig = SamplerConfig(
topK = 50,
topP = 0.95,
temperature = 0.8
)
)
)

Send the prompt

LiteRT-LM

conversation.sendMessageAsync(prompt)
.collect { message -> Log.d(message.toString()) }
conversation.close()
engine.close()

What to use?

It depends 😌

BOTTOM LINE · WHAT TO USE?Gemini Nano≈ macOSVia ML Kit GenAI and Prompt APIOne model, chosen by the OSNo bring-your-own modelGuard railsLimited availabilityOEM adoption is key for its futureLiteRT-LM≈ Arch LinuxCompletely self managedAny model you can convertWider availability

Size still matters

Caveats

  • •
    Don't expect SOTA performance
  • •
    Focus on a narrow use case
  • •
    Start from a fine-tuned model
  • •
    Handling resource usage is challenging

Shipping it at iFood

A real on-device AI feature in production

Meet iFood

A Brazilian food delivery app

  • •
    Users order food and groceries through the app
    • •200 M+ customers
    • •120 M+ orders/month
  • •
    After checkout, they follow the order in real time
  • •
    That experience is what we call Waiting

The waiting screen

A high-attention moment after checkout.

The iFood Waiting screen after checkout

A simple user problem

What if the order is not for the user?

  • •
    Someone else may be waiting for the order
  • •
    They need status, ETA, address or pickup details
  • •
    Today, that conversation usually happens on WhatsApp

Version 1

A floating AI button at the bottom of the waiting screen.

Why it didn't ship

The model wasn't competing with another model. The feature was competing with the product.

  • •
    The bottom area already had an important product surface
  • •
    The FAB introduced a new competing interaction
  • •
    More clicks here could mean fewer clicks somewhere else
  • •
    The safest decision was not to launch it

Version 02

One button with one job. A much quieter place in the product.

Version 2: Share in the toolbar

Generating on-device

The loading state is actual local inference not a request waiting for our backend.

Completing the job

Success isn't generating text. It's helping the user send it.

Small models need constraints

A technically valid answer can still be completely wrong for the product.

  • •
    Greeted the restaurant instead of the recipient
  • •
    Mixed delivery vocabulary with pickup orders
  • •
    Returned internal status terminology
  • •
    Invented details when the promo left room for interpretation

Make the model boring

We don't want creativity here. We want a faithful transformation of known facts.

MAKE THE MODEL BORINGOrder factsrestaurantstatusaddresstime windowpickup codeRules & vocabularyRemote ConfigDelivery vs pickup vocabularyUse only known factsNever invent missing informationOn-device modelWhatsApp messageYour order from Pizza Place is ready for pickup.123 Main St, 5–5:30 PM.Pickup code: 4827.✓✓

Can we show the button?

The feature only exists when inference can run right now.

CAN WE SHOW THE BUTTON?Featureenabled?YESDevicesupported?YESModelavailable?NOHideNOHideUNAVAILABLEHideDOWNLOADABLEBackground downloadMaybe next orderAVAILABLE✓Show Share

Coverage is a product metric

Before asking who clicked the feature, we need to know who could even see it.

  • •
    Device support
  • •
    Model availability
  • •
    Background download state
  • •
    Remote config audience
  • •
    Waiting version eligibility

Measure before rollout

At iFood scale, even limited device coverage means thousands of opportunities every day.

  • •
    ~2.3% coverage
  • •
    100K+ eligible orders/day
IFOOD COVERAGETechnically eligible orders per day020K40K60K80K100K120KEligible orders / day106KJun 28104KJul 0596KJul 12103KJul 19109KJul 26116KAug 02115KAug 09116KAug 16Source: iFood production data · Gabriel Volpi

What happened in production?

Local inference worked. But it wasn't instant.

  • •
    8.71 s median generation time
IFOOD LATENCYOn-device generation latency8.08.59.09.510.010.511.0123456789101112Generation time (seconds)P90 · 9.61 sP50 · 8.71 sSource: iFood production data · Gabriel Volpi

What we learned

Most of the hard problems weren't model problems.

  • •
    Start with one user job
  • •
    Placement matters as much as capability
  • •
    Hide is a valid fallback
  • •
    Constrain small models aggressively
  • •
    Measure coverage, not only inference

From demo to product

On-device AI is a product decision, not an SDK demo.

FROM DEMO TO PRODUCTModelPromptProductPlacementCoverageMetricsProduction

Press X to doubt

This is a fad.

L.A. Noire

L.A. Noire

It's a movie, not a picture

3 ideas

  • •
    Race to the bottom
  • •
    History
  • •
    Specialization + ubiquity
  • •
    Keep an eye on the trend lines
RACE TO THE BOTTOM · ON-DEVICE LLMS, 1 YEAR AGO VS TODAY1 year agoTodayModel size> 8.0 GB2.6 GB (Gemma 4 E2B)Prefill speed< 10 tokens / s3,808 tokens / sOutput speed< 3 tokens / s52 tokens / sFine-tuning methodFull model fine-tuneParameter-efficient fine-tuningon top of a shared foundation modelQuantizationHW accelerationEfficient fine-tuningSource: LiteRT-LM, Gemma 4 E2B on Galaxy S26 Ultra (GPU)

It's a movie, not a picture

Remember Mainframes?

Mainframe, 1970s-80s

Mainframe, 1970s-80s

It's a movie, not a picture

Remember Mainframes?

Terminal

Terminal

It's a movie, not a picture

Remember Mainframes?

Early personal computer

Early personal computer

Specialization and ubiquity

Mini models everywhere

  • •
    Paper: Small Language Models are the Future of Agentic AI
  • •
    “For many agentic workloads, SLMs are a superior default to LLMs due to cost, latency, controllability, and fine-tuning ease.”
  • •
    Specialized small models, inside systems that also use bigger ones

Cheaper inference, new business models

Specialization and ubiquity

  • •
    Free apps
  • •
    Premium app tiers
  • •
    Single payment apps
  • •
    No variable cost per user interaction with LLMs

Small models everywhere

GBoard Writing Tools

GBoard Writing Tools

Small models everywhere

Pixel Screenshots

Pixel Screenshots

Give it a go

google-ai-edge/gallery

google-ai-edge/gallery

Give it a go

AI Edge Gallery

AI Edge Gallery

Give it a go

DenisovAV/flutter_gemma

DenisovAV/flutter_gemma

Give it a go

flutter_gemma sample app

flutter_gemma sample app

References

Putting the genie in the bottle

  • •
    ML Kit GenAI APIs
  • •
    LiteRT-LM
  • •
    Gemma 4
  • •
    Paper: Small Language Models are the Future of Agentic AI
  • •
    Deep Dive into LLMs like ChatGPT
  • •
    Google AI Edge

Thank you!

Iury Souza
QR code for iurysouza.dev
@iurysza
iurysouza.dev
Gabriel Volpi
@_gvolpi
@iFood

Questions?

Iury Souza
QR code for iurysouza.dev
@iurysza
iurysouza.dev
Gabriel Volpi
@_gvolpi
@iFood