All Insights
#PromptInjection
August 2026

Sunday Coffee & Code: turning my Google Pixel phone into an LLM server

A week or ago ago Brandon Laur (The White Hatter) mentioned he was curious about the local language models I have been running on my phone in LM Playground. That offhand comment sent me down a rabbit hole, and this weeks SC&C is what came out the other end.

By Steve Harris

๐—ฆ๐˜‚๐—ป๐—ฑ๐—ฎ๐˜† ๐—–๐—ผ๐—ณ๐—ณ๐—ฒ๐—ฒ & ๐—–๐—ผ๐—ฑ๐—ฒ: ๐˜๐˜‚๐—ฟ๐—ป๐—ถ๐—ป๐—ด ๐—บ๐˜† ๐—š๐—ผ๐—ผ๐—ด๐—น๐—ฒ ๐—ฃ๐—ถ๐˜…๐—ฒ๐—น ๐—ฝ๐—ต๐—ผ๐—ป๐—ฒ ๐—ถ๐—ป๐˜๐—ผ ๐—ฎ๐—ป ๐—Ÿ๐—Ÿ๐—  ๐˜€๐—ฒ๐—ฟ๐˜ƒ๐—ฒ๐—ฟ

A week or ago ago Brandon Laur (The White Hatter) mentioned he was curious about the local language models Iโ€™ve been running on my phone in LM Playground. That offhand comment sent me down a rabbit hole, and this weekโ€™s SC&C is what came out the other end.

LM Playground is a neat Android app that runs GGUF models fully on-device and offline. Download a model, load it, chat. No cloud, no API keys. Great for privacy, and a genuinely interesting sandbox for the sort of small models I keep writing about.

But I wanted more than a chat window. If the model is already running on the phone, ๐—ฐ๐—ผ๐˜‚๐—น๐—ฑ ๐—œ ๐—ฟ๐—ฒ๐—ฎ๐—ฐ๐—ต ๐—ถ๐˜ ๐—ณ๐—ฟ๐—ผ๐—บ ๐—ผ๐—ป๐—ฒ ๐—ผ๐—ณ ๐—บ๐˜† ๐˜€๐˜†๐˜€๐˜๐—ฒ๐—บ๐˜€ ๐—ฎ๐˜ ๐—ต๐—ผ๐—บ๐—ฒ? So I forked the repo and sat down with Claude Code to add an OpenAI-compatible API server that binds to the phoneโ€™s WiFi interface. ๐—ง๐—ต๐—ฒ ๐—ถ๐—ฑ๐—ฒ๐—ฎ: expose the loaded model at the standard /v1/chat/completions endpoint, protected by a generated bearer key, reachable from any other device on the same network.

A morning of pair-programming later, it works and ๐—บ๐˜† ๐—ฝ๐—ต๐—ผ๐—ป๐—ฒ ๐—ถ๐˜€ ๐—ป๐—ผ๐˜„ ๐—ฎ ๐—น๐—ถ๐˜๐˜๐—น๐—ฒ ๐—ถ๐—ป๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐˜€๐—ฒ๐—ฟ๐˜ƒ๐—ฒ๐—ฟ. I can point the openai Python SDK at http://:8080/v1 and get completions back from a model running on a Pixel 8.

OK, what do you do with a pocket-sized LLM endpoint? Test it, of course - Iโ€™ve wired it up to my prompt-injection testing harness, and as I write this itโ€™s grinding through a 500-prompt run, throwing injection attempts at the on-device model and logging how it responds. Maybe the kind of thing Brandonโ€™s world cares about: if people are going to run private models on their phones, how robust are they?

The build itself was a proper yak-shave. Learned a lot about Android Studio, NDK, CMake version pins, a Vulkan backend that wanted a shader compiler I didnโ€™t have. I learned more about the ggml build system than I strictly meant, or wanted to. In the end I turned Vulkan off and the build went clean.

All in all, a perfect Sunday C&C project, fun, learned something new, did some more testing (๐˜ข๐˜ฏ๐˜ฅ ๐˜ข๐˜ญ๐˜ด๐˜ฐ ๐˜Š๐˜ญ๐˜ข๐˜ถ๐˜ฅ๐˜ฆ ๐˜Š๐˜ฐ๐˜ฅ๐˜ฆ ๐˜ฅ๐˜ฐ๐˜ฆ๐˜ด ๐˜ข ๐˜ฃ๐˜ณ๐˜ช๐˜ญ๐˜ญ๐˜ช๐˜ข๐˜ฏ๐˜ต ๐˜ซ๐˜ฐ๐˜ฃ ๐˜ฐ๐˜ฏ ๐˜ˆ๐˜ฏ๐˜ฅ๐˜ณ๐˜ฐ๐˜ช๐˜ฅ ๐˜ข๐˜ฑ๐˜ฑ๐˜ด) and also that on-device models are more accessible than most people realise

Iโ€™ll share the results of the injection run in my prompt injection testing repo. For now the coffeeโ€™s gone cold and the testโ€™s still running.

๐—ฅ๐—ฒ๐—ฝ๐—ผโ€˜๐˜€

Want to Discuss This Topic?

Steve is always happy to have a direct conversation.