# Running Large Language Models Fully Offline on Mobile with React Native > Paolo Pecis explores fully offline speech-to-text and LLM summarization on iOS and Android with React Native, whisper.rn, llama.rn, and the SoloAI app. - URL: https://lokahq.github.io/tech-blog/running-large-language-models-fully-offline-on-mobile-with-react-native/ - Type: Blog article - Authors: Paolo Pecis (Senior Mobile Developer) - Published: 2026-10-05 - Reading time: 10 min - Tags: React Native, LLM, Machine Learning, Mobile Development, Offline AI --- In the last few years, running Machine Learning models locally has become increasingly viable, even on constrained devices like smartphones. With the right tradeoffs, it’s now possible to run **speech-to-text**, **text generation**, and other ML workloads directly on a mobile device without relying on cloud APIs. This approach enables **offline-first** experiences and **privacy-by-design** applications, where sensitive data is processed entirely on the user’s device. In this article, I’ll walk through my research on running **local LLMs on mobile** using **[React Native](https://reactnative.dev/)** (a cross platform development framework based on TypeScript), the libraries I evaluated, the limitations I encountered, and how all of this materialized into a real showcase app called **SoloAI**. ## Why Local ML on Mobile? Running models locally has a few major advantages: - **Offline-first** experiences - **User privacy**, no data sent to third parties - **Lower operational costs**, no inference servers - **Lower latency**, especially for short tasks The obvious downside is that mobile devices have limited CPU, memory, storage, and battery. The challenge is finding the right balance between model size, performance, accuracy, and user experience. ## Evaluated Libraries ### TensorFlow Lite via [react-native-fast-tflite](https://github.com/mrousavy/react-native-fast-tflite) I initially explored **react-native-fast-tflite**, a performant React Native bridge for TensorFlow Lite models. I tested some models sourced from [Kaggle](https://www.kaggle.com/models?tfhub-redirect=true): - [Spam detection](https://www.kaggle.com/models/tensorflow/spam-detection) classifiers - [MobileBERT](https://www.kaggle.com/models/tensorflow/mobilebert) for question answering - [EfficientDet](https://www.kaggle.com/models/tensorflow/efficientdet) for object detection using the device camera [![Spam detection with TensorFlow Lite](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/simulator_screenshot_2CA4774F-AB90-4E8D-ACA3-66C5845BCB2F.webp)](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/simulator_screenshot_2CA4774F-AB90-4E8D-ACA3-66C5845BCB2F.webp) Spam detection with TensorFlow Lite [![Question answering with MobileBERT](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/simulator_screenshot_E8736F1D-C238-4DA3-8B4D-8F0A17D29A53.webp)](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/simulator_screenshot_E8736F1D-C238-4DA3-8B4D-8F0A17D29A53.webp) Question answering with MobileBERT [![Object detection with EfficientDet on Android](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Screenshot_20260128_152545_LokaNativeMobileLLM.jpg)](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Screenshot_20260128_152545_LokaNativeMobileLLM.jpg) Object detection with EfficientDet on Android This setup worked, but the results for visual object detection were inconsistent and not always accurate enough for a polished product. Performance also varied significantly across devices. Moreover, these TFLite models are also harder to manage compared to the Whisper and LLaMA-based approaches described later because each model needs to be analyzed and checked: which inputs and outputs it requires, how to parse tokens, and various other constraints. The preparation code is quite extensive: ```jsx import { loadTensorflowModel } from 'react-native-fast-tflite'; //Load the model from local path const handleLoadModel = async () => { setLoading(true); try { const loadedModel = await loadTensorflowModel(mobileBertModel); setModel(loadedModel); setModelInfo(modelToString(loadedModel)); } catch (e) { Alert.alert('Error', 'Failed to load MobileBERT model'); } finally { setLoading(false); } }; //Run and parse the model const handleRun = useCallback(async () => { if (!model || !question.trim() || !passage.trim()) { Alert.alert('Error', 'Please provide both question and passage'); return; } try { const inputs = prepareInputs(question, passage); // MobileBERT typically expects 3 inputs: input_ids, attention_mask, token_type_ids const modelInputs = [ inputs.inputIds, inputs.attentionMask, inputs.tokenTypeIds, ]; const results = await model.run(modelInputs); const result = processModelOutputs(results, inputs); setOutput(result); } catch (error) { Alert.alert('Error', `Failed to run model: ${error}`); setOutput(null); } }, [model, question, passage]); ``` ```tsx //Helper only for MobileBERT input/output mapper import RAW_VOCAB from '@assets/vocab.json'; const VOCAB = RAW_VOCAB as Record; export const MAX_SEQ_LENGTH = 384; // Standard for BERT models export const MAX_QUERY_LENGTH = 64; export interface QAResult { answer: string; confidence: number; startIndex: number; endIndex: number; } export interface ModelInputs { inputIds: Int32Array; attentionMask: Int32Array; tokenTypeIds: Int32Array; passageStartIndex: number; passageTokens: number[]; } // Build a reverse lookup map once const REVERSE_VOCAB: Record = Object.entries(VOCAB).reduce( (acc, [token, id]) => { acc[id] = token; return acc; }, {} as Record, ); // 2) Tokenizer: split on whitespace and map to IDs via your full VOCAB export const tokenize = (text: string): number[] => { // lower-case, strip punctuation (keep only letters/numbers/space), then split const cleaned = text.toLowerCase().replace(/[^\w\s]/g, ''); // <-- remove punctuation const words = cleaned.split(/\s+/); return words.map(w => VOCAB[w] ?? VOCAB['[UNK]']); }; // 3) Prepare BERT inputs exactly as before export const prepareInputs = ( question: string, passage: string, ): ModelInputs => { const CLS = VOCAB['[CLS]']; const SEP = VOCAB['[SEP]']; const PAD = VOCAB['[PAD]']; // tokenize & truncate question const qTokens = tokenize(question).slice(0, MAX_QUERY_LENGTH); let pTokens = tokenize(passage); // limit passage length so total ≤ MAX_SEQ_LENGTH const maxPassageLen = MAX_SEQ_LENGTH - qTokens.length - 3; if (pTokens.length > maxPassageLen) { pTokens = pTokens.slice(0, maxPassageLen); } // build [CLS] + question + [SEP] + passage + [SEP] const inputIdsArr: number[] = [CLS, ...qTokens, SEP, ...pTokens, SEP]; // pad up to MAX_SEQ_LENGTH while (inputIdsArr.length < MAX_SEQ_LENGTH) { inputIdsArr.push(PAD); } // attention mask = 1 for real tokens, 0 for PAD const attentionMaskArr = inputIdsArr.map(id => (id === PAD ? 0 : 1)); // token type IDs: 0 for question segment, 1 for passage segment const tokenTypeIdsArr = new Array(MAX_SEQ_LENGTH).fill(0); const passageStartIndex = qTokens.length + 2; // after [CLS]+question+[SEP] for ( let i = passageStartIndex; i < passageStartIndex + pTokens.length + 1; i++ ) { tokenTypeIdsArr[i] = 1; } return { inputIds: Int32Array.from(inputIdsArr), attentionMask: Int32Array.from(attentionMaskArr), tokenTypeIds: Int32Array.from(tokenTypeIdsArr), passageStartIndex, passageTokens: pTokens, }; }; // 4) Post‐process logits back into text answer export const processModelOutputs = ( results: any[], inputs: ModelInputs, ): QAResult => { if (results.length < 2) { throw new Error('Expected at least 2 outputs (start_logits, end_logits)'); } const startLogits = results[0] as Float32Array; const endLogits = results[1] as Float32Array; let bestScore = -Infinity; let bestStart = inputs.passageStartIndex; let bestEnd = inputs.passageStartIndex; const passageEnd = inputs.passageStartIndex + inputs.passageTokens.length; const maxAnswerLen = 30; // find best span in the passage region for (let i = inputs.passageStartIndex; i < passageEnd; i++) { for (let j = i; j < Math.min(i + maxAnswerLen, passageEnd); j++) { const score = startLogits[i] + endLogits[j]; if (score > bestScore) { bestScore = score; bestStart = i; bestEnd = j; } } } // extract the token IDs for that span (relative to passage) const startOffset = bestStart - inputs.passageStartIndex; const endOffset = bestEnd - inputs.passageStartIndex; const answerIds = inputs.passageTokens.slice(startOffset, endOffset + 1); // map back to tokens and join const answerTokens = answerIds.map(id => REVERSE_VOCAB[id] ?? '[UNK]'); const answerText = answerTokens.join(' ').replace(/##/g, ''); // approximate confidence via sigmoid of bestScore const confidence = 1 / (1 + Math.exp(-bestScore)); return { answer: answerText || 'No answer found', confidence, startIndex: bestStart, endIndex: bestEnd, }; }; // Default sample data export const DEFAULT_PASSAGE = 'France is a country in Western Europe. Paris is the capital and largest city of France. The city has a population of over 2 million people and is known for landmarks like the Eiffel Tower and the Louvre Museum.'; export const DEFAULT_QUESTION = 'What is the capital of France?'; ``` This complexity pushed me toward focusing on **text-based and audio-based use cases**, where the results were far more reliable and the setup easier. ### Speech to Text with [whisper.rn](https://github.com/mybigday/whisper.rn) **whisper.rn** is a React Native bridge around whisper.cpp, a highly optimized C++ implementation of OpenAI’s Whisper models. It worked well, providing good transcription accuracy, even with small models, fully offline inference, and predictable performance on modern smartphones. For the MVP, I used the [ggml-tiny.en.bin](https://huggingface.co/ggerganov/whisper.cpp/blob/main/ggml-tiny.en.bin) model, which is around **77 MB** and small enough to be bundled directly with the app. ```jsx import { initWhisper, TranscribeFileOptions, TranscribeResult, WhisperContext, } from 'whisper.rn/index.js'; //Initialize the whisper model useEffect(() => { const initializeWhisper = async () => { try { whisperContextRef.current = await initWhisper({ filePath: require('../../../../assets/models/ggml-tiny.en.bin'), }); setIsModelLoaded(true); } catch (error) { console.error('Error initializing Whisper model:', error); } }; initializeWhisper(); }, []); //Run the model and get the transcritpion of the audioResoursce const handleTranscribeAudio = useCallback(async (audioResource: string) => { setIsTranscribing(true); const options: TranscribeFileOptions = { language: 'en', prompt: 'Your name is SoloAI, and you are an helpful assistant that transcribes audio files to store them as notes.', }; if (!whisperContextRef.current) { Alert.alert('Error', 'Whisper model is not loaded.'); return; } try { const { promise } = whisperContextRef.current.transcribe( audioResource, options, ); const result = await promise; setTranscription(result.result); } catch (error) { console.error('Error transcribing audio:', error); const errorMessage = error instanceof Error ? error.message : 'Failed to transcribe audio.'; Alert.alert('Error', errorMessage); } finally { setIsTranscribing(false); } }, []); ``` ### Text Generation with [llama.rn](https://github.com/mybigday/llama.rn) For text generation and summarization, I used **llama.rn**, a React Native binding of llama.cpp. This enables fully offline LLM inference and support for GGUF models (LLaMA, Mistral, etc.) with no dependency on cloud APIs. ```jsx import { initLlama, LlamaContext, RNLlamaOAICompatibleMessage } from 'llama.rn'; //Initialize the llama model const initializeLlama = useCallback(async () => { setLlamaInitialized(false); try { if (llamaModelPath) { const context = await initLlama({ model: llamaModelPath, use_mlock: true, n_ctx: 2048, n_gpu_layers: 1, }); llamaContextRef.current = context; setLlamaInitialized(true); } } catch (error) { console.error('Error initializing Llama model:', error); setLlamaInitialized(false); } }, [llamaModelPath]); //Run the model and get the summary of the giver transcription const handleSummarization = useCallback(async (transcription: string) => { if (!llamaContextRef.current || !transcription) { console.warn('Llama context or transcription is not available'); return; } try { setIsSummarizing(true); const messages: RNLlamaOAICompatibleMessage[] = [ { role: 'system', content: LLAMA_SUMMARIZATION_PROMPT, }, { role: 'user', content: transcription, }, ]; const result = await llamaContextRef.current?.completion({ messages: [...messages], n_predict: 100, stop: [...stopWords], }); return result.text; } catch (error) { console.error('Error during summarization:', error); } finally { setIsSummarizing(false); } }, []); ``` The model used in the MVP is: [Llama-3.2-1B-Instruct-IQ4\_XS\_imat.gguf](https://huggingface.co/medmekk/Llama-3.2-1B-Instruct.GGUF/blob/main/Llama-3.2-1B-Instruct-IQ4_XS_imat.gguf) This model is approximately **743 MB**, which introduces an important constraint. ## Handling Large Models on Mobile Bundling a 700+ MB model directly inside a mobile app bundle is not realistic because of app store size limits, long install times and the overall poor user experience. ### The Compromise SoloAI uses a **manual model download** approach: - The app works offline by default. - **The only time an internet connection is required** is when the user downloads the LLaMA model. - Once downloaded, the model is stored locally and reused indefinitely. This setup keeps the app app-store friendly, transparent about storage usage and fully offline after initial setup. [![One-time model download](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_17_Pro_Max_-_2026-01-28_at_15.57.14.webp)](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_17_Pro_Max_-_2026-01-28_at_15.57.14.webp) One-time model download ## Downloading the LLaMA Model with [react-native-fs](https://github.com/itinance/react-native-fs) To handle large file downloads, I used **react-native-fs**, which provides reliable filesystem access on both iOS and Android. Below is a simplified example of how the LLaMA model is downloaded and stored locally. ### Example: Model Download Function ```jsx import RNFS from 'react-native-fs'; const LLAMA_MODEL_URL = 'https://your-cdn-or-storage.com/Llama-3.2-1B-Instruct-IQ4_XS_imat.gguf'; const MODEL_DESTINATION_PATH = `${RNFS.DocumentDirectoryPath}/llama.gguf`; export const downloadLlamaModel = async ( onProgress?: (progress: number) => void, ): Promise => { const download = RNFS.downloadFile({ fromUrl: LLAMA_MODEL_URL, toFile: MODEL_DESTINATION_PATH, progress: (res) => { const progressPercent = (res.bytesWritten / res.contentLength) * 100; onProgress?.(progressPercent); }, progressDivider: 1, }); const result = await download.promise; if (result.statusCode !== 200) { throw new Error('Failed to download LLaMA model'); } return MODEL_DESTINATION_PATH; }; ``` Once downloaded, this path is passed directly to **llama.rn** when initializing the model context. After this step, the app can run **fully offline**, including transcription and summarization. ## Technology Stack Overview The entire project is built using **React Native**, targeting both iOS and Android from a single TypeScript codebase. The core ML stack is composed of: - **[whisper.rn](https://github.com/mybigday/whisper.rn)**: offline speech-to-text - **[llama.rn](https://github.com/mybigday/llama.rn)**: offline text generation and summarization - **[@simform\_solutions/react-native-audio-waveform](https://github.com/SimformSolutionsPvtLtd/react-native-audio-waveform)**: audio recording with waveform visualization. A custom patch was applied to support **uncompressed .wav format**, required by Whisper (the library originally supported only .m4a) - [@op-engineering/op-sqlite](https://github.com/OP-Engineering/op-sqlite): Used to store all data locally using SQLite - **[Redux](https://redux.js.org/)**: state management ## App Showcase ### Onboarding & Login - Carousel onboarding shown on first app launch - Explains core features and offline benefits - Secure access via biometric authentication or PIN - Ensures notes and data remain private [![Offline onboarding and security setup](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_10.58.30.webp)](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_10.58.30.webp) Offline onboarding and security setup [![PIN and biometric login](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_10.58.44.webp)](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_10.58.44.webp) PIN and biometric login ### Main Dashboard - List of all notes with: - Category - Duration - Date - Title - First lines of transcription - Search bar with filtering and sorting to quickly find notes [![Recordings dashboard and search](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_11.59.06.webp)](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_11.59.06.webp) Recordings dashboard and search [![Recording filters and sorting](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_11.19.33.webp)](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_11.19.33.webp) Recording filters and sorting --- ### Record Screen - Start and stop voice recording - Real-time audio waveform visualization - After recording: - Modal to choose category - Editable note title [![Voice recording with a live waveform](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_10.33.26.webp)](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_10.33.26.webp) Voice recording with a live waveform [![Recording title and category selection](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_10.46.45.webp)](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_10.46.45.webp) Recording title and category selection --- ### Note Detail Screen - Full transcription view - Generated summary - Actions: - Edit title - Change category - Copy transcription or summary - Delete note [![Locally generated transcription](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_10.49.16.webp)](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_10.49.16.webp) Locally generated transcription [![Locally generated summary](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_12.04.58.webp)](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_12.04.58.webp) Locally generated summary [![Changing a note’s category](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_12.05.02.webp)](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_12.05.02.webp) Changing a note’s category --- ### Settings & Category Management - Change PIN code - Terms & Conditions - Contact information - Delete all notes - Manage categories: - Create new categories - Edit name and color - Delete existing categories [![SoloAI settings](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_11.35.20.webp)](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_11.35.20.webp) SoloAI settings [![Managing note categories](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_11.35.24.webp)](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_11.35.24.webp) Managing note categories [![Editing category names and colors](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_11.35.29.webp)](https://lokahq.github.io/tech-blog/blog/running-large-language-models-fully-offline-on-mobile-with-react-native/Simulator_Screenshot_-_iPhone_16_Pro_-_2026-01-28_at_11.35.29.webp) Editing category names and colors ## Conclusion This research demonstrates that **running ML models locally on mobile devices is not only feasible, but practical**. While there are clear limitations related to performance, precision, and model size, the results show strong potential for fully offline applications, privacy-first user experiences and reduced dependency on cloud infrastructure. SoloAI serves as a concrete example of how local ML can be used to build meaningful, user-facing products that respect user privacy and work even without an internet connection. You can download and test SoloAI App Store: [SoloAI by Loka](https://apps.apple.com/us/app/soloai-by-loka/id6756312615) Play Store: [SoloAI by Loka](https://play.google.com/store/apps/details?id=com.loka.lokamobileaitranscribe)