# Building RAG Systems on AWS: Lessons from Serverless and EC2 Benchmarks > A practical comparison of RAG ingestion and retrieval architectures using in-memory vector stores across AWS Lambda, EC2, and local environments. - URL: https://lokahq.github.io/tech-blog/building-rag-systems-on-aws-lessons-from-serverless-and-ec2-benchmarks/ - Type: Blog article - Authors: Romulo Pagnozzi (Senior ML Engineer), Crhistian Cardona (ML Engineering Manager) - Published: 2026-01-20 - Updated: 2026-09-30 - Reading time: 22 min - Tags: AWS, RAG, Vector Database, Benchmarking, Serverless - Topics: inference, benchmarks - GitHub: https://github.com/LokaHQ/vector-db-benchmark - Originally published on Medium: https://medium.com/loka-engineering/building-rag-systems-on-aws-lessons-from-serverless-and-ec2-benchmarks-165b481a0c95 --- ### Written by: Romulo Pagnozzi — Senior ML Engineer, Crhistian Cardona — ML Engineering Manager This blog post explores the evolving landscape of Retrieval-Augmented Generation (RAG) systems through the lens of server-based and serverless, AWS-native architectures, with a particular focus on in-memory vector databases running on AWS Lambda, classical EC2 instances, and the emergence of Amazon S3 Vectors as a new paradigm for scalable, cost-efficient vector storage and retrieval. We begin by framing the rise of Generative AI (GenAI) and the continued relevance of RAG for overcoming context window, cost, and latency constraints in modern Large Language Models (LLMs). As semantic retrieval becomes a core building block of GenAI applications, the choice of vector database is increasingly critical — especially under real-world constraints around privacy, scalability, and cost efficiency. To guide these architectural decisions, we introduce a modular and extensible benchmarking framework that evaluates vector databases under realistic ingestion and inference workflows — both locally and on AWS EC2, with selected experiments running on AWS Lambda. The framework focuses on production-oriented metrics, including memory consumption, latency, cold-start overhead, and storage footprint. A central theme of the post is the evolution of Amazon S3 Vectors — from its public preview launch on July 15 to its global availability announced at AWS re:Invent in December 2025. We analyze how S3 Vectors introduces a fundamentally different, object-storage-native approach to vector search, enabling durable, scalable, and low-operations RAG architectures that align naturally with serverless execution models. We also examine supported deployment patterns, including ephemeral in-memory vector stores with S3-based persistence, and compare custom RAG pipelines against managed alternatives such as AWS Bedrock Converse, highlighting trade-offs in flexibility, cost, and control. To ground the analysis in practical relevance, the post presents a business-oriented use-case matrix that maps real-world RAG applications to their performance, security, and compliance requirements. Finally, we introduce the vector technologies included in the benchmark — Chroma, LanceDB, Milvus Lite, Qdrant, and Amazon S3 Vectors — and outline future plans to expand coverage to widely used libraries such as Faiss, Annoy, and Redis Vector, further enabling data-driven decisions for teams building GenAI systems on AWS. ### TL;DR — Key Takeaways 1. **Compute dominates RAG performance.**\ Embedding generation and LLM inference accounts for most latency and cost; vector storage and retrieval add bounded, predictable overhead across both serverless and EC2 environments. 2. **Serverless RAG is production-ready.**\ Ephemeral, in-memory vector stores running on AWS Lambda can deliver reliable, low-latency ingestion and search, making them well suited for event-driven and elastic workloads with minimal operational burden. 3. **Architecture matters more than the database.**\ There is no universally “best” vector database — outcomes depend on execution context and workload patterns. AWS-native, object-store-backed designs (such as S3-based persistence) enable scalable, low-ops architectures that are especially compelling for cost-sensitive and privacy-constrained systems. --- ## RAG Still Matters Nowadays, generative AI (GenAI) has become one of the most exciting and disruptive areas in the technology landscape. Far beyond simple automation, GenAI is unlocking new forms of reasoning, decision-making, and creativity — largely enabled by breakthroughs in Large Language Models (LLMs) like GPT-5, Claude 4.5, DeepSeek among others. These models are capable of generating high-quality human-like text, answering complex questions, and synthesizing insights from large corpora of data. In the early stages of LLM development, however, one key limitation was the restricted context window — typically a few hundred or thousand tokens. This meant that only a small portion of information could be considered at a time, making it challenging to reference long documents or support complex multi-turn conversations. To work around this, a new architectural pattern emerged: Retrieval-Augmented Generation (RAG). RAG systems address the context limitation by retrieving relevant information from an external knowledge base and injecting that information directly into the prompt of the LLM at runtime. This allows models to access dynamic, task-specific, and up-to-date information without retraining. Crucially, retrieval is based on semantic similarity rather than keyword search, enabling much more intelligent and flexible lookups. As RAG gained popularity, a range of implementation strategies evolved. One core component of RAG is the chunking strategy, which defines how input documents are broken into smaller segments before indexing. Chunking is important because too-large chunks may dilute relevance, while too-small chunks may lack meaningful context. Some common strategies include: - **Fixed-size chunking**, where text is split into equal-length character or token blocks. - **Recursive chunking**, which preserves semantic units (like paragraphs or sentences) while trying to meet a target size. - S**emantic chunking**, which uses NLP models to split text at natural topic boundaries. Another critical element is the embedding system, which converts text chunks into high-dimensional vectors that can be compared based on cosine or dot-product similarity. These embeddings act as a numeric “fingerprint” of the content. Common embedding providers include: - OpenAI embeddings (text-embedding-3-small and others) - Amazon Bedrock (e.g., Titan Embeddings) - Cohere and HuggingFace models - Sentence-transformers, for open-source options with multi-lingual support Over time, a wide variety of RAG architectures have emerged, each addressing different retrieval and reasoning needs: - **Standard RAG** uses simple retrieval with chunk injection into the prompt. - **Graph RAG** represents documents or concepts as nodes and edges, allowing for multi-hop or relationship-aware retrieval. - **Multi-hop RAG** retrieves across several linked documents before synthesizing an answer. - **Agentic RAG** combines retrieval with reasoning agents that perform multi-step planning, filtering, and synthesis. Even though today’s LLMs can process millions of tokens in a single context window, RAG remains highly relevant. Why? Because large context windows come at a significant cost in latency and price, and are often wasteful for focused use cases. RAG offers a more scalable and cost-efficient approach, especially for applications like document search, chat over private data, or API response augmentation. ## A Spectrum of Vector Database Architectures In this post, we benchmark and compare five vector database solutions: Chroma, LanceDB, Milvus Lite, Qdrant, and Amazon S3 Vectors. These were selected based on their relevance to RAG (Retrieval-Augmented Generation) applications, especially in serverless and cost-efficient architectures. Each tool represents a different philosophy of vector database design — from lightweight in-memory libraries to cloud-native, fully managed solutions. Before diving into the benchmarks, let’s briefly introduce each system and its typical use case. See **Table 1** for a high-level feature comparison. [Embedded content](https://datawrapper.dwcdn.net/0LRqG/1/) **Table 1**: Key features of In-Memory Vector Databases Evaluated in this Benchmark ### Chroma Chroma is an open-source, lightweight vector database purpose-built for LLM-native workflows. It was designed with simplicity in mind — offering easy-to-use Python APIs, tight integration with popular embedding providers, and automatic metadata handling. - **History**: Launched in 2023 with the rise of LangChain and RAG tooling, Chroma quickly became popular in the open-source community for prototyping LLM applications. - **Architecture**: In-memory by default, but can persist indexes to local disk (SQLite-based). - **Use Case:** Ideal for small-scale LLM applications, fast prototyping, and serverless environments (like AWS Lambda) with fast load times. ### LanceDB LanceDB is a vector database built on top of the Lance columnar data format, which itself was designed by contributors to the Apache Arrow ecosystem. It combines the efficiency of memory-mapped file access with the structure of SQL-style queries. - **History**: Created by the team at eto.ai (formerly part of the Apache Arrow project) to enable efficient vector analytics over large datasets. - **Architecture**: Uses disk-based, columnar storage with optional memory mapping for acceleration. Compatible with S3 and DuckDB. - **Use Case**: Best for analytics-heavy or hybrid workflows involving both structured and unstructured data, or when vectors need to be accessed from object storage in a scalable way. ### Milvus Lite Milvus Lite is a minimal version of the full Milvus vector database, designed for local or in-memory vector search. Under the hood, it uses Faiss for indexing and search. - **History**: Released by Zilliz, the creators of Milvus, to support developers who needed Milvus’ speed and indexing capabilities without the overhead of a full server. - **Architecture**: Runs in-memory, no persistence unless manually implemented; very lightweight and simple to deploy. - **Use Case**: Suitable for quick demos, testing, and small-scale applications, especially when you need Faiss-level speed in a self-contained script or notebook. ### Qdrant Qdrant is a full-featured, production-grade vector database written in Rust. It offers robust performance, filtering, metadata handling, and strong API support (REST, gRPC, and SDKs in multiple languages). - **History**: Developed by the Qdrant team in 2021, with a focus on being developer-friendly and high-performance, it quickly gained traction in open-source and enterprise settings. - **Architecture**: Uses RocksDB for persistence and supports several index types like HNSW and PQ. - **Use Case:** Ideal for large-scale, production vector search, multi-tenant applications, or when tight control over filtering and metadata is needed. ### Amazon S3 Vectors Amazon S3 Vectors is a new cloud-native vector store built directly into Amazon S3. It allows users to store, index, and search vector embeddings without managing any infrastructure or dedicated database engines. ![S3 Vector architecture. Source : Amazon blog](https://lokahq.github.io/tech-blog/blog/building-rag-systems-on-aws-lessons-from-serverless-and-ec2-benchmarks/0-wbBXkz7EHcyjiQ1E.webp) **Figure 1: S3 Vector architecture. Source**: Amazon blog - **History**: Announced in July 2025 at the AWS Summit New York as a public preview, this service represents Amazon’s push toward serverless, cost-efficient vector search at scale. Its general availability was later announced at AWS re:Invent 2025, marking a major milestone in AWS’s vector-native offerings. - **Architecture**: Fully managed; vectors are stored in S3 and accessed via dedicated vector APIs. Ideal for infrequent but scalable queries over large corpora. **Use Case**: Suited for cold storage retrieval, long-term embedding storage, and RAG applications where storage cost (overall 90% less) and durability matter more than ultra-low latency. ## General Use Case The goal of this study is to evaluate selected in-memory vector databases in the context of a realistic and modern use case — one where data privacy, security, and cost-efficiency are top priorities. In particular, we are focused on a serverless implementation that avoids persistent infrastructure, reduces overhead, and enables low-cost, on-demand retrieval of vector data for AI-enhanced applications. This setup is especially relevant in enterprise environments where sensitive data must be processed securely, such as internal document search, legal review, healthcare assistance, or customer support. In these scenarios, embedding and storing data in memory — rather than in centralized long-lived services — can reduce the attack surface, ensure per-session isolation, and improve control over lifecycle management. To help ground the benchmarking in practical relevance, we’ve defined a series of business-oriented use cases where these tools could be applied. Each use case highlights not only the application characteristics but also the related security considerations that influence vector DB selection. [Embedded content](https://datawrapper.dwcdn.net/iFb8U/1/) **Table 2**: Business use cases This catalog of use cases shows the diversity of scenarios where in-memory vector search with optional persistence is not only viable but often preferable. These use cases also shape our benchmarking criteria: they demand fast startup, isolated memory usage, minimal persistence, and strong alignment with security principles — all within a serverless or lightweight computer architecture. In the next section, we will describe the architecture for the benchmarking experiment and define the metrics used to evaluate each tool. ## Structure of the Benchmarking Framework The benchmarking framework we’ve developed is designed to be modular, extensible, and vendor-agnostic — capable of testing a wide variety of vector databases regardless of their architecture or underlying technology. Whether you’re working with an in-memory library like Chroma or a fully managed service like Amazon S3 Vectors, the tool adapts to support your environment and constraints. Currently, our primary focus is on serverless architectures, particularly AWS Lambda, where we evaluate: - **In-memory vector libraries** - **Disk-based vector stores** - **Cloud object store-based solutions** This extensible design allows us to benchmark the full RAG (Retrieval-Augmented Generation) pipeline — both locally and in AWS Lambda environments — and to simulate production-ready execution patterns. ![: AWS reference architecture for AWS Lambda/EC2/ECS applications.](https://lokahq.github.io/tech-blog/blog/building-rag-systems-on-aws-lessons-from-serverless-and-ec2-benchmarks/1-AmQ_e-sACB2gr6quSF2DVw.webp) **Figure 2**: AWS reference architecture for AWS Lambda/EC2/ECS applications. Our current emphasis is on serverless architectures, particularly AWS Lambda, where speed, cost, and security are critical. In parallel, we run the same experiments on EC2 instances (and extensible to ECS), since some in-memory vector databases cannot be executed reliably within Lambda’s constraints. Within this setup, we evaluate three classes of storage approaches — in-memory vector libraries, disk-based vector stores, and cloud object storage solutions — to understand how each behaves under ephemeral execution models. This methodology enables us to benchmark the entire RAG pipeline in both local and AWS Lambda environments, closely simulating production-ready execution patterns. **Figure 2** illustrates the serverless architecture, with Lambda functions conceptually replaceable by EC2 or ECS components for non-serverless deployments. The framework follows three core design principles. First, it uses reusable workflows with a plugin‑like structure, making it easy to swap out embedding models, chunking strategies, and vector databases. Second, it applies structured benchmarking, running systematic experiments across configurations to provide reproducible data for comparison. Finally, it is aligned with production use, so the same pipelines tested here can be deployed directly into real systems without major rework. Together, these principles reduce risk, speed up delivery, and give decision‑makers confidence based on real metrics rather than assumptions. To capture realistic performance, the benchmarking breaks the RAG lifecycle into two main workflows: document upload and search/inference. The document upload workflow measures how quickly an uploaded file can be parsed, chunked, embedded, and stored. Documents are first parsed (e.g., from PDFs), then split into semantically meaningful chunks, transformed into embeddings, inserted into an in‑memory vector store, and finally persisted to a backend such as S3 or EFS. Metrics such as processing time, RAM and CPU usage, and storage size are collected to identify bottlenecks. The goal is to optimize how quickly and cost‑effectively documents can be prepared for search. The search and inference workflow measures how fast and accurately the system can respond to a user query. In this phase, the relevant vector index is loaded from storage into memory, the user’s query is embedded, and a similarity search retrieves the most relevant chunks. These chunks are then formatted into a prompt and sent to the LLM for answer generation. Here, we track metrics like vector index load time, search latency, LLM response latency, and total inference cost to ensure the system can deliver high‑quality results at scale. In some scenarios, we also test bypassing retrieval entirely and submitting the full document directly to the LLM to compare performance with and without RAG. Under the hood, the framework supports a wide range of configuration parameters. It can run inside Lambda or containers, handle different document sizes and types, and work with multiple LLM providers such as Titan, OpenAI, or Cohere. Storage backends are pluggable, supported by design S3, EFS, and EBS. This flexibility makes it possible to evaluate many realistic deployment patterns without rewriting the benchmarking logic. While ChromaDB and Qdrant support true in-memory operation suitable for AWS Lambda’s serverless environment, both Milvus Lite and LanceDB have architectural limitations that prevented them from running in Lambda’s constrained runtime. Consequently, we deployed these benchmarks on a local machine to accommodate their requirements. However, the experimental conditions remained consistent across all vector databases: each system processed the same corpus of documents, fetched data from S3, persisted the constructed vector database back to S3 for storage, and subsequently loaded the persisted vectors from S3 to perform inference queries. This approach ensured that despite the different hosting environments, all databases underwent identical data ingestion, persistence, and retrieval workflows, maintaining the validity of performance comparisons. A key feature of the architecture is its security‑oriented design. Rather than relying on a centralized, long‑lived vector database, the system creates an ephemeral, in‑memory store on the fly for each Lambda invocation, scoped to a single document or user session. This approach minimizes the attack surface, gives full control over data lifecycle and deletion, and avoids persisting sensitive information unless explicitly required — a major advantage for regulated environments such as healthcare or legal services. Finally, we collect detailed performance metrics tailored to both EC2 instances and Lambda’s execution model. These include memory and CPU requirements to estimate cost per execution, I/O performance for reading and writing index files from S3 or EFS, startup overhead when loading vector indexes into memory, and end-to-end latency from request to response. We also measure cold-start sensitivity to assess whether performance degrades significantly on first invocation. The central question driving this benchmarking effort is straightforward: can an ephemeral, serverless, in-memory vector store architecture deliver the performance, security, and scalability required for real-world RAG applications? And how does it compare with equivalent EC2 and ECS-based implementations using the same in-memory databases? By systematically evaluating a diverse set of vector database technologies under identical conditions, this framework is designed to produce the data-driven insights needed to make these architectural decisions with confidence. For more details, you can check the github repository: [_**https://github.com/LokaHQ/vector-db-benchmark/**_](https://github.com/LokaHQ/vector-db-benchmark/) ## Inside the Benchmarking Engine To rigorously evaluate vector database performance for Retrieval-Augmented Generation (RAG) workflows, we built a modular benchmarking engine that mimics a production-grade pipeline — but with hooks for observability, customization, and post-analysis. The entire system is designed using well-known software patterns like Pipeline, Strategy, Factory, and Adapter, which make it both extensible and maintainable. This modular architecture lets us plug in any combination of embedding models, vector databases, chunking strategies, and storage backends — without rewriting core logic. Let’s walk through the main components of this engine. Each run tests combinations of documents, embedding models, and vector DBs. Performance is tracked at every stage: ```php text, text_extraction_time = time_operation("Extracting text", lambda: extract_text_from_doc(...)) ``` ### Pluggable Vector Databases Each vector DB (Chroma, Qdrant, LanceDB, etc.) implements a shared interface. A factory dynamically loads the right backend: ```ini db_class = get_database_provider(job.vector_db_name) ``` This allows clean benchmarking across memory, disk, and cloud-native DBs like Amazon S3 Vector. ### Embedding Providers Local models (e.g., MiniLM) and cloud-based (e.g., Bedrock Titan) share the same logic: ```ini embedder = create_embedder(model_name) embeddings = embedder.embed_texts(texts) ``` Async support ensures scalability even with large batches and high concurrency. ### S3-Based Persistence To support Lambda, the system stores vector DBs as tarballs in S3 — downloaded and loaded into memory per request: ```scss aws_client.upload_directory_as_tar(local_dir, bucket, key) ``` This enables stateless, secure, serverless RAG workflows. ### Structured Metrics Benchmark results (latency, RAM, CPU) are tracked in Pydantic models and visualized via pandas: ```lua df.groupby(["vector_db_type"])[["total_time"]].mean() ``` This helps compare cost-performance trade-offs between tools. ## Alternatives to RAG Systems in AWS When building GenAI-powered applications on AWS, one of the early design decisions involves choosing how to handle document retrieval and context injection. Do you build a custom Retrieval-Augmented Generation (RAG) pipeline from scratch, or leverage AWS Bedrock’s more streamlined Converse API? Each option has its strengths — and understanding the trade-offs is key to choosing the right fit for your use case. The Converse API offers a simple, powerful interface for conversational applications. It abstracts away many of the complexities involved in traditional RAG systems — such as document chunking, vector embedding, and similarity search — allowing developers to upload documents and ask questions about them in a single API call. For teams looking to move fast and avoid managing infrastructure, this approach is very appealing. However, there are some important limitations. At the time of writing, the Converse API allows a maximum of five documents per conversation, with each document capped at 4.5 MB. Once the limit is reached, your application must either block further uploads or replace existing files. These constraints can work well for many scenarios, but they may be restrictive in use cases involving large documents or dynamic file uploads across a session. Pricing is another important consideration. Bedrock pricing is based on input and output tokens — and when using documents with the Converse function, their contents are included in the input token count. A standard prompt may use only a few hundred tokens, but a single 4.5 MB document can contribute 100,000 to 200,000 tokens, depending on its content. That means a session with two documents and a prompt could easily exceed 200,000 tokens per request. At current pricing, this could result in a cost of roughly $0.30–$0.60 per interaction — reasonable for high-value use cases, but something to monitor in high-volume or consumer-facing applications. For these reasons, many teams turn to custom RAG implementations when they need more flexibility or control. A custom pipeline enables selective chunking of documents, caching of embeddings, fine-tuned vector searches, and tailored cost optimizations. You can also support much larger files and document collections than Converse currently allows. ## Results Our benchmarking follows the end-to-end lifecycle of a Retrieval-Augmented Generation (RAG) system, covering ingestion, search, and consolidated performance analysis. During ingestion, we measured how each vector database performs under serverless and EC2-based constraints, tracking document size scaling, chunk creation, embedding generation, insertion time, and resource usage (CPU and memory). By breaking ingestion into sub-steps, we identified which stages dominate overall cost as data volume increases. In the search phase, we evaluated query-time behavior, including database load time, query embedding latency, vector search time, and total retrieval latency. This allowed us to separate cold-start effects from actual similarity search performance. Finally, we combined ingestion and search metrics to analyze trade-offs holistically, reflecting real-world RAG usage where storage, performance, and cost must be evaluated together rather than in isolation. ### Ingestion Time vs Document Size The EC2-based execution exhibits a more stable and predictable performance profile, with text extraction and embedding times scaling smoothly as document size increases and consistently low overhead across all size buckets. This behavior highlights the advantages of a long-lived environment, where resources remain warm and readily available, avoiding repeated initialization costs. In contrast, the Lambda-based execution with an in-memory database follows a similar overall scaling trend but shows higher relative overhead, especially for very small and very large files. These additional costs are primarily driven by cold starts, runtime setup, and the stricter CPU and memory constraints of ephemeral executions. Notably, the core pipeline stages — text extraction, embedding, and vector insertion — remain consistent across both setups, confirming that serverless ingestion does not alter the fundamental workload characteristics, but rather introduces a fixed execution overhead. In practice, this makes EC2 a strong fit for sustained, batch-oriented ingestion, while Lambda stands out as a cost-efficient and operationally simpler option for event-driven or intermittent ingestion workloads, without compromising correctness or scalability. ![EC2 — Ingestion Pipeline Time Breakdown by File Size](https://lokahq.github.io/tech-blog/blog/building-rag-systems-on-aws-lessons-from-serverless-and-ec2-benchmarks/1-d0SMIxo-Fw_9f1t72YvUGQ.webp) **Figure 3: **EC2 — Ingestion Pipeline Time Breakdown by File Size ## Ingestion: Performance Overview In the EC2-based execution, the ingestion benchmark exhibits a more stable and clearly differentiated performance profile across vector databases. Ingestion time scales predictably with document size, and resource utilization — particularly peak RAM and CPU — remains consistent for each database, reflecting the benefits of a long-lived, provisioned environment. Disk-backed and hybrid databases such as Qdrant and LanceDB show slightly higher but steady memory footprints, while Chroma’s in-memory design results in higher peak RAM usage that grows with both document size and chunk count. Importantly, the ingestion time breakdown reveals that text extraction and embedding dominate total latency across all databases, with minimal overhead variability, making EC2 well suited for sustained, high-throughput ingestion where predictability and throughput are priorities. ![a: EC2 RAG ingestion Bechmark — Performance Overview (by Vector DB)](https://lokahq.github.io/tech-blog/blog/building-rag-systems-on-aws-lessons-from-serverless-and-ec2-benchmarks/1-h0KVsiVNpDQkPzC1pL4czA.webp) **Figure 4-a: **EC2** **RAG ingestion Bechmark — Performance Overview (by Vector DB) ![b: Serverless RAG ingestion Bechmark — Performance Overview (by Vector DB)](https://lokahq.github.io/tech-blog/blog/building-rag-systems-on-aws-lessons-from-serverless-and-ec2-benchmarks/1-rUzOLF5p_yd_nZmbr9QXQQ.webp) **Figure 4-b: **Serverless RAG ingestion Bechmark — Performance Overview (by Vector DB) In contrast, the Lambda-based execution with in-memory databases preserves the same overall scaling trends but introduces noticeably higher dispersion in ingestion time, CPU utilization, and overhead — especially for smaller documents. This variability is driven by cold starts, repeated runtime initialization, and stricter resource constraints inherent to ephemeral serverless execution. While core stages such as text extraction, embedding, and vector insertion behave similarly in both environments, Lambda amplifies fixed costs relative to workload size, making overhead more visible. Persisted database size and chunk-related scaling remain consistent between EC2 and Lambda, confirming that the underlying data model and indexing behavior are unchanged. In practice, this highlights a clear trade-off: EC2 excels for continuous, batch-oriented ingestion with tight performance control, whereas Lambda offers a simpler, cost-efficient option for event-driven or sporadic ingestion workloads — accepting some variability without compromising correctness or scalability. ### Ingestion Pipeline Breakdown Both the stacked and file-size–bucketed breakdowns reveal a clear and consistent pattern across execution models and vector databases: - **Text extraction** remains relatively stable across document sizes and contributes only a small fraction of the total ingestion time. - **Embedding generation** is the dominant cost and scales directly with document size and chunk count, quickly becoming the primary driver of ingestion latency. - **Vector insertion and persistence** increase with the number of chunks but remain secondary compared to the cost of embedding inference. These results confirm that embedding computation — not storage or indexing — is the primary ingestion bottleneck. Even when using S3-backed vector storage, persistence overhead remains modest relative to compute-heavy stages. In practice, this means that optimizing ingestion performance is far more dependent on embedding model choice, batching strategies, and compute allocation than on the underlying vector storage backend. ## Search: Performance Overview In the EC2-based execution, the search benchmark exhibits a stable and highly predictable performance profile across all vector databases. Database load time and total search latency show a clear separation of concerns: once the vector index is loaded, search time remains consistently low and largely independent of top-k values. Vector search latency scales smoothly, and the distribution plots highlight tight variance, particularly for Qdrant, Lance, and Milvus Lite, which benefit from persistent memory, warm caches, and uninterrupted CPU availability. Importantly, LLM completion latency remains largely orthogonal to the vector database choice, confirming that retrieval performance does not interfere with downstream generation once the search step is complete. ![a: EC2 RAG Benchmark — By Vector DB](https://lokahq.github.io/tech-blog/blog/building-rag-systems-on-aws-lessons-from-serverless-and-ec2-benchmarks/1-vJ09mn1NrMFwxinJ53r5nw.webp) **Figure 5-a: **EC2 RAG Benchmark — By Vector DB ![b: Serverless RAG Benchmark — By Vector DB](https://lokahq.github.io/tech-blog/blog/building-rag-systems-on-aws-lessons-from-serverless-and-ec2-benchmarks/1-7sTO1rTJznTuaUlR_oehRA.webp) **Figure 5-b: **Serverless RAG Benchmark — By Vector DB In contrast, the Lambda-based execution with in-memory vector databases follows the same qualitative trends but introduces additional variability and overhead. Search total time becomes more sensitive to database load time, especially for solutions that require index hydration at invocation time. This effect is most visible in the wider latency distributions and higher tail latencies, driven by cold starts, memory reallocation, and ephemeral storage constraints. However, the vector search operation itself remains fast and well-bounded across databases, indicating that Lambda does not alter the computational complexity of retrieval — only the surrounding execution context. In practice, this reinforces a clear trade-off: EC2 is better suited for latency-critical, high-throughput search workloads, while Lambda provides a compelling, cost-efficient option for bursty or event-driven RAG queries where slight increases in variance are acceptable without compromising correctness or scalability. ### Search Pipeline Breakdown The search-phase benchmarks reveal a similarly consistent pattern across vector databases and execution environments: - **Vector search time** is generally low and stable for in-memory databases, even as `top_k` increases, indicating efficient ANN indexing and query execution. - **Database load time** becomes a key differentiator in serverless setups, contributing a noticeable fixed overhead — particularly for in-memory databases that must be rehydrated on each invocation. - **LLM completion latency** dominates end-to-end query time and shows the strongest correlation with input token count rather than vector database choice. Overall, the results show that retrieval itself is not the primary bottleneck in RAG search pipelines. Instead, latency is driven by a combination of database initialization (in ephemeral environments) and downstream LLM inference. Importantly, S3-backed vector storage introduces higher per-query retrieval latency but avoids repeated loading costs, while in-memory databases deliver faster raw search performance at the expense of cold-start overhead. In practice, this highlights a trade-off between raw query speed and operational simplicity, reinforcing that search performance is shaped more by execution context and LLM behavior than by the vector index alone. ## End-to-End RAG Implications Taken together, the ingestion and search benchmarks point to a clear architectural takeaway: - **Compute-heavy stages (embedding generation and LLM inference) dominate both latency and cost** across all environments. - **Vector storage and retrieval introduce predictable, bounded overhead**, which is secondary compared to model execution — regardless of whether vectors are stored in memory or in S3-backed systems. As a result, vector databases are rarely the limiting factor in real-world RAG pipelines. Teams can safely adopt serverless and object-store-backed vector architectures, such as Amazon S3 Vectors, without compromising correctness or scalability. The main trade-off shifts from raw performance to operational characteristics: in-memory stores favor ultra-low latency for warm, high-throughput workloads, while S3-backed approaches optimize for durability, cost efficiency, and simplicity in event-driven or elastic systems. Ultimately, optimizing RAG performance means focusing on model choice, batching strategies, and execution context — rather than over-indexing on vector storage mechanics. ## Final Thoughts This benchmarking study demonstrates that serverless RAG architectures — when combined with in-memory vector stores and persistent, object-store-backed solutions like Amazon S3 Vectors — form a robust, scalable, and cost-efficient alternative to traditional always-on vector databases. Across ingestion and search workloads, our evaluation of Chroma, LanceDB, Milvus Lite, and Qdrant shows that even ephemeral, per-request vector stores can deliver consistent, production-ready performance within the operational constraints of AWS Lambda. A key outcome of the results is that compute, not storage, dominates end-to-end RAG performance. Embedding generation and LLM inference account for the majority of latency and cost, while vector indexing, persistence, and retrieval introduce relatively small and predictable overhead. This holds true even when vectors are stored in S3-backed systems, confirming that object storage does not materially degrade RAG performance for typical workloads. Within this landscape, Amazon S3 Vectors emerges as a compelling architectural option. By combining the durability, scalability, and cost model of S3 with native vector indexing and search, it enables semantic retrieval without the operational burden of managing dedicated vector infrastructure. For teams prioritizing cost control, data isolation, and deployment simplicity — especially in event-driven or elastic workloads — this approach meaningfully simplifies RAG system design on AWS. From the benchmarks, three practical recommendations stand out. For serverless RAG workloads on AWS, Amazon S3 Vectors is the best overall choice, offering a strong balance of performance, scalability, and operational simplicity. For local development and rapid prototyping, Chroma remains an excellent fit, thanks to its lightweight design and fast iteration cycle. For long-running, high-QPS, latency-sensitive systems, Qdrant continues to excel, providing mature features and consistently low retrieval latency in always-on environments. In the end, there is no universal solution. The right vector database depends on workload patterns, performance expectations, and operational constraints. What these results make clear, however, is that AWS-native, serverless RAG architectures are no longer a compromise — they are increasingly the simplest, most scalable, and most cost-efficient way to build real-world RAG systems. The future of RAG on AWS is lean, elastic, and unmistakably vectorized. ## Future Work: Expanding Benchmark Coverage As we continue to evolve our benchmarking framework, our next steps focus on broadening its flexibility and testing capabilities across a wider variety of storage and deployment scenarios. First, we plan to add parametric support for multiple storage backends — including Amazon EFS, EBS, and ephemeral disks — to measure how intermediate disk processing affects performance. In parallel, we’ll extend testing beyond AWS Lambda to include containerized deployments on Kubernetes and Amazon ECS, enabling apples‑to‑apples comparisons between serverless and container-based architectures. We’re also adding wrappers for several widely-used vector libraries and databases, which will allow us to benchmark a broader spectrum of performance profiles and deployment patterns: - **HNSWlib** — Lightweight, high-speed approximate nearest neighbor (ANN) search for resource-constrained or embedded environments. - **Annoy** — A read-only, disk-backed index optimized for minimal memory usage and fast load times — ideal for serverless scenarios. - **Faiss** — Meta’s high-performance vector library with GPU support and advanced indexing options, widely used in large-scale systems. - **Chroma** — A Python-native vector database tailored for LLM applications, popular for rapid prototyping and small-scale RAG pipelines. - **Milvus Lite** — A standalone, minimal version of Milvus for quick testing and ephemeral workloads. - **Redis Vector** — Adds vector search capabilities to Redis, combining in-memory speed with built-in persistence. These additions will make our framework far more versatile and allow deeper comparisons across cost, latency, memory usage, and suitability for serverless, on-device, and production-scale deployments — giving teams the data they need to choose the right architecture with confidence. ## References - - - [https://www.trychroma.com](https://www.trychroma.com/) - - - [https://qdrant.tech](https://qdrant.tech/)