
Running Large Language Models Fully Offline on Mobile with React Native
Paolo Pecis explores fully offline speech-to-text and LLM summarization on iOS and Android with React Native, whisper.rn, llama.rn, and the SoloAI app.
From the people building it

Paolo Pecis explores fully offline speech-to-text and LLM summarization on iOS and Android with React Native, whisper.rn, llama.rn, and the SoloAI app.
Authentication and authorization for AI agents on AWS. Who is allowed to call your agent, and how your agent proves itself to everything it calls.


Paolo Pecis explores fully offline speech-to-text and LLM summarization on iOS and Android with React Native, whisper.rn, llama.rn, and the SoloAI app.

Authentication and authorization for AI agents on AWS. Who is allowed to call your agent, and how your agent proves itself to everything it calls.

A native PyTorch benchmark showing EvolutionaryScale ESMC-300M running on AWS Trainium2: 490.8 variants per second on a single trn2.3xlarge — approximately 42.4 million protein variants scored per day at approximately $1.26 per million, 60% below H100 in this fixed-shape D2Deep benchmark.

Everything is buildable now. Every team I meet lives in the same climate: more ideas than quarters to test them in, more tools than problems, roadmaps that read like wish lists because nothing on them is technically…

The main post covers what we found: Grok 4.3 lands in the same 96–97% accuracy band on GSM8K whether you reach it through Amazon Bedrock Mantle or the first-party xAI API, and low reasoning effort captures most of the…

A public-facing benchmark story on Grok 4.3 through Amazon Bedrock Mantle and the xAI API, focused on accuracy, latency, cost, reliability, and enterprise deployment tradeoffs.

If you work anywhere near data-heavy science — genomics, metagenomics, proteomics — you’ve probably hit the same wall: building reliable, reproducible data pipelines is hard. Nextflow is an open-source workflow engine…

Modern computational drug discovery depends on complex workflows that connect AI models, structural biology tools, databases, and quality-control steps into reproducible pipelines. While many organizations have access…

Training a vision model on millions of images is, more than anything, a data problem wearing an ML costume. The architecture can many times be the least interesting part. Everything around it is usually more fun: how we…

In the digital audio subset of AI, the questions currently on everyone’s mind are

When building AI agents with AWS Bedrock, guardrails are your first line of defense for keeping conversations on topic. A common requirement is to restrict an agent to a pre-defined strict allowlist, blocking everything…

Fourteen open-source text-to-speech models across six emotion-control paradigms, benchmarked against AWS Nova Sonic v2 on naturalness, expressiveness, latency, and licensing.

What to know before you scale cofolding jobs from five structures to ten thousand.

How we built an agent that ports Python ML code to edge hardware for space.

The main post covers what we found: GPT-5.5 on Amazon Bedrock matches OpenAI API accuracy and comes out faster on every latency and throughput dimension we measured. This post covers how we ran it. If you want to…

A comparison of GPT-5.5 through Amazon Bedrock and the OpenAI API across answer quality, latency, throughput, and reliability, followed by two production-shaped copilot builds.

A practical benchmark showing Hugging Face Bio Carbon running on AWS Trainium2 with NxD Inference, covering what worked, what we measured, and why it matters for bio and HCLS teams moving open models into production.

Hugging Face Bio’s Carbon release is interesting for two reasons at once. First, it is a biology model. Second, it is not an infrastructure outlier. Many genomic models come with custom architectures, specialized…

Lineage for a CDC Pipeline on AWS

How to turn scattered guidelines into adaptive workflows that score, explain, and improve content at scale

A Portuguese playwright wanted to push AI beyond the assistant role and onto the theatre stage. Here’s how LOKA built a live system that brings a cast of AI actors to perform alongside humans.

Productionizing Agentic Use Cases in Weeks, Not Months

If you work in regulated domains (Healthcare, Life Sciences, Finance) you routinely hit constraints that break the default “just call a hosted API” approach:

An ablation-driven study of GRPO-style reinforcement learning with LoRA on Trinity Mini for biomedical relation extraction, covering the training choices that materially changed performance.

When building stable and maintainable software, one of the biggest challenges developers face is that modern applications often need to support multiple versions of a feature—different algorithms or formats that can…

A practical benchmark of RAG ingestion and search performance across AWS Lambda and EC2

Solving the Challenges of Traditional Pipelines

After exploring the foundations of Agentic Patterns in Part 1, where we looked at more structured workflows like Sequential and Parallel workflows, LLM Routing, and Reflection, this time we are going to push further…

These days, Agents are everywhere in daily conversation and it seems everyone wants one. We read or hear about AI Agents, Agentic AI, Agentic Architectures, and Agentic Patterns almost every day. But what do these terms…

Loka’s participation in the recent nf-core hackathon in Medellín, Colombia, culminated in the successful development of MITNANEX, a pipeline designed to extract and assemble mitochondrial genomes, identify and annotate…

Moving from XML to Jetpack Compose helps developers write cleaner, faster and more modern UI code — essential for staying up to date with Android’s best practices.

Simple interface for data stored in S3

A DeepSeek benchmark on AWS that compares cost, performance, and deployment choices for business workloads using open-weight models.

A deployment walkthrough for serving distilled DeepSeek-R1 models with Amazon SageMaker AI, including the AWS setup for open-weight reasoning models.
Engineering notes on deploying DeepSeek-R1 models to AWS-designed silicon instances and serving open-weight reasoning models on AWS.

Deploy distilled DeepSeek-R1 open-weight models through Amazon Bedrock, with a hands-on AWS setup from the Loka engineering team.

How the open-recipe LLM is transforming GenAI

You probably have heard the history of how HTTP was created and you might even know about the new HTTP/3 version. In this article, I want to focus on some of the lesser-known details of HTTP and why we need a new…