# Running Six Open-Source Cofolding Models on AWS: Lessons Learned from a Compute Perspective > A compute benchmark of six open-source cofolding models on AWS, measuring runtime, memory, and estimated cost across 29 PROTAC ternary complexes. - URL: https://lokahq.github.io/tech-blog/running-six-open-source-cofolding-models-on-aws-lessons-learned-from-a-compute-perspective/ - Type: Blog article - Authors: Julián F. Fernández (Bioengineering Lead), Andres Florian (Bioinformatics Engineer), Daniel Sabogal (Junior Bioengineer) - Published: 2026-06-30 - Updated: 2026-09-30 - Reading time: 12 min - Tags: AlphaFold, Drug Discovery, AI, Biology, Chemistry - Topics: biomedical-ai, benchmarks, inference - GitHub: https://github.com/jwohlwend/boltz, https://github.com/chaidiscovery/chai-lab, https://github.com/bytedance/Protenix, https://github.com/RosettaCommons/foundry, https://github.com/aqlaboratory/openfold-3, https://github.com/Biohub/esm - Datasets: https://www.nature.com/articles/s41467-025-61272-5 - Originally published on Medium: https://medium.com/loka-engineering/running-six-open-source-cofolding-models-on-aws-lessons-learned-from-a-compute-perspective-d34715c60a5b --- _Co-authored by_ [_Andrés Florian_](https://www.linkedin.com/in/andres-felipe-f-45631b223) _and_ [_Daniel Sabogal_](https://www.linkedin.com/in/daasabogalro/)_._ What to know before you scale cofolding jobs from five structures to ten thousand. Running modern cofolding models is no longer just a question of choosing the GPU with the most VRAM. In practice, different models stress different parts of the system: GPU inference, CPU-side preprocessing, system RAM, and orchestration overhead. The model with the smaller VRAM footprint can easily be the one that ties up an instance the longest. Models such as [AlphaFold3](https://www.nature.com/articles/s41586-024-07487-w), [Boltz-2](https://github.com/jwohlwend/boltz), and [RoseTTAFold3](https://github.com/RosettaCommons/foundry) can now predict multi-chain biomolecular complexes (proteins, nucleic acids, and small molecules) with impressive accuracy. But published accuracy benchmarks rarely answer a practical question: _**how much does it actually cost to run them, and what hardware do you need?**_ GPU compute is expensive and scarce: a mid-range cloud GPU instance (NVIDIA L4) costs close to $1 per hour on AWS, and top-tier accelerators (H100, L40S) command several times more, so picking the wrong model/instance combination wastes both money and time. To get concrete answers, we benchmarked six open-source cofolding models (Boltz-2, [Chai-1](https://github.com/chaidiscovery/chai-lab), [Protenix](https://github.com/bytedance/Protenix), RoseTTAFold3, [OpenFold3](https://github.com/aqlaboratory/openfold-3), and [ESMFold2](https://github.com/Biohub/esm)) on a single demanding workload: PROTAC ternary-complex prediction. We measured compute, runtime, and cost behavior on AWS, not biological accuracy (that is covered in a companion analysis). The question throughout is the one you face before committing budget: _**which model and instance combinations are actually efficient before you scale up?**_ ## The Benchmark Setup Our workload was deliberately demanding. PROTACs (PROteolysis TArgeting Chimeras) are bifunctional molecules that recruit an E3 ubiquitin ligase to degrade a target protein, a clinically compelling modality capable of reaching targets that conventional small-molecule inhibitors cannot. The resulting _ternary complex_ (two proteins bridged by a flexible chemical linker) is a tough compute stress test: multiple chains, a flexible bifunctional small molecule, and non-trivial MSA inputs. For the benchmark set we used 22 solved ternary structures from [TernaryDB](https://www.nature.com/articles/s41467-025-61272-5), plus 7 PROTAC structures released to the PDB after the documented training cutoffs of all six models: 10AY (2026–04–22), 9Q03 (2025–10–15), 9SAF (2025–10–01), 9SAI (2025–10–29), 9YA9 (2026–02–18), 9SV3 (2025–12–03), and 9T32 (2025–12–03). The recent cases include large three-chain complexes (9SV3 at 816 aa, 9Q03 and 9YA9 at 753 aa), which widen the input-size range to 256–816 total amino acids and make the runtime-vs-size analysis below credible across a real spread. | Model | Upstream source | Version / commit | | ------------ | ----------------------- | -------------------------------- | | Boltz-2 | jwohlwend/boltz | boltz==2.2.1 (tag v2.2.1) | | Chai-1 | chaidiscovery/chai-lab | tag v0.6.1 | | ESMFold2 | Biohub/esm | commit c94ed8d | | OpenFold3 | OpenFold3 (build clone) | commit c9bfe23, ckpt of3-p2-155k | | Protenix | bytedance/Protenix | protenix==2.0.0, ckpt v1.0.0 | | RoseTTAFold3 | RosettaCommons/foundry | commit cd4d94f (RF3) | **Table 1: **Model versions behind each container. Checkpoints as run. Each of the six models was containerized in a dedicated Docker image and deployed as a SageMaker Processing Job (on \`ml.\`-prefixed instances, e.g. \`ml.g6.2xlarge\`), running one structure at a time with a single random seed. Model versions are listed in Table 1. MSAs were pre-generated outside the benchmark and provided from S3 storage, so MSA search time is excluded from all reported runtimes. Every model ran with its published default hyperparameters, except that we normalised the number of diffusion samples to five per run for every model so the comparison is sample-matched (RoseTTAFold3 emits six: five plus a ranked best). We did no per-case tuning and did not otherwise equalise internal workloads (recycling steps, MSA depth) across architectures. The primary comparison ran all six models on a single instance type, \`g6.2xlarge\` (NVIDIA L4 GPU, 24 GB VRAM, 8 vCPUs, 32 GB RAM), for an apples-to-apples comparison. To see how GPU tier affects performance, we additionally ran Boltz-2 across four \`2xlarge\` instance types (8 vCPU held fixed, so the GPU is the main variable) spanning the T4 to L40S range (Table 2); those instance-level comparisons used a 5-case subset and should be read directionally. | Instance | GPU | VRAM | vCPU | $/hr | | ------------ | ---- | ----: | ---: | ----: | | g4dn.2xlarge | T4 | 16 GB | 8 | $0.75 | | g5.2xlarge | A10G | 24 GB | 8 | $1.21 | | g6.2xlarge | L4 | 24 GB | 8 | $0.98 | | g6e.2xlarge | L40S | 48 GB | 8 | $2.24 | **Table 2: **Instance types used (provisioned via SageMaker Processing Jobs as the corresponding `ml.` instances). Listed prices (from June 2026) are approximate raw EC2 on-demand rates for `us-east-1` (Linux); SageMaker bills a premium over these, so they are lower bounds (see the cost discussion) ## Two Observed Runtime Profiles The most striking result from profiling these jobs is that not all models use compute in the same way. The figure below shows median model-compute time and peak memory footprint (VRAM and system RAM) for all six models on the \`g6.2xlarge\` instance. ![Median model-compute time and peak GPU and system memory across six cofolding models on g6.2xlarge for 29 PROTAC complexes.](https://lokahq.github.io/tech-blog/blog/running-six-open-source-cofolding-models-on-aws-lessons-learned-from-a-compute-perspective/1-5WQniJ_Cylt8AvKyCs8aMw.webp) **Figure 1: **Median model-compute time (Panel A) and peak memory footprint (Panel B) for six cofolding models on `g6.2xlarge` (NVIDIA L4, 24 GB VRAM), across 29 PROTAC ternary complexes (5 diffusion samples per run). Panel B shows peak VRAM and peak system RAM, the two limits that gate which instance a model fits on; the dotted line marks the 16 GB tier. System RAM, not VRAM, is the binding constraint for OpenFold3 (26 GB). Hatched bars in Panel A mark models with a preprocessing-heavy profile. **Inference-dominant profile.** In our runs, Boltz-2, Protenix, RoseTTAFold3, and Chai-1 behaved as inference-dominant workloads in this benchmark. Mean GPU utilization ranges from 37% to 75%, median model-compute time falls between 74 and 424 seconds, and peak system RAM stays at or below 12.5 GB. Wall-clock time largely tracked GPU inference; preprocessing was fast relative to the forward pass in our setup. **Preprocessing-heavy profile.** ESMFold2 and OpenFold3 showed a different pattern. Mean GPU utilization across the full job is only 13% and 10%, respectively, not because the GPU is unused, but because GPU inference occupied only a fraction of total wall-clock time. In our runs, the remainder appeared to be spent on CPU/RAM-side work (OpenFold3 in particular ran the highest mean CPU load, about 22%, versus 13–15% for the others). Depending on the model, this may include MSA parsing and featurization, pair representation construction, ligand featurization, or other input-preparation steps. These stages likely contributed substantially to wall-clock time and, in our setup, did not translate into sustained GPU utilization, although we did not perform stage-level tracing. We therefore describe this as an observed deployment profile rather than an intrinsic architectural property. The first surprise was that VRAM was not always the limiting factor. For OpenFold3, the hard constraint was system RAM: it required over 26 GB in our runs, ruling out 16 GB instances, while ESMFold2 peaked at 18 GB (Panel B of the figure above and the RAM column of Table 3). For these two models, instance selection should consider CPU performance and system RAM, not only GPU tier. For OpenFold3 in particular, upgrading the GPU alone was not the main driver of performance in our instance comparison. **Sample count is a memory knob.** Normalizing every model to five diffusion samples exposed a second memory surprise. Going from one to five samples barely changed Boltz-2’s peak VRAM (≈10 GB) but roughly doubled its model-compute time (61 s to 126 s); for ESMFold2 the opposite held, with runtime essentially flat but peak VRAM growing enough to exceed the 24 GB L4 on most cases. ESMFold2 ran within the L4 for only 14 of 29 cases; 12 required a 48 GB L40S, and three (9Q03, 9SV3, 9YA9, the largest three-chain complexes) ran out of memory even on the L40S. How a model spends additional samples, in time or in memory, is therefore itself a deployment property: more samples can quietly change which instance you need, not just how long you wait. ## Does a better GPU pay off? The second practical question was whether a larger GPU actually paid for itself. For inference-dominant workloads, upgrading the GPU tier seems like a natural way to reduce runtime and cost. We tested this with Boltz-2, the model for which we have the strongest evidence of inference-dominant behavior: we ran it across multiple GPU tiers and observed meaningful runtime differences across instances. We used the same five TernaryDB cases and the same Boltz-2 container (the default single-sample build; this study isolates the GPU tier, and the speedup ratio is independent of sample count) on four \`2xlarge\` SageMaker instance types spanning T4 to L40S. Runtime below is split into model-compute time (the model process itself) and billed time (model compute + the near-fixed container/data overhead SageMaker also charges for); we lead with model-compute time because production would amortize the overhead by batching (see the cost section). The model container, inputs, and inference settings were held fixed; the instance type was varied. Note that CPU generation, vCPU count, memory, and I/O characteristics may vary across instance families, so the GPU is not the only difference between these instances. The Figure 2 and Table 3 below show the results. Given the small case count (five structures), these results are indicative rather than definitive. ![Boltz-2 across four 2xlarge GPU tiers (median over 5 cases, 8 vCPU held fixed so the GPU is the main variable).](https://lokahq.github.io/tech-blog/blog/running-six-open-source-cofolding-models-on-aws-lessons-learned-from-a-compute-perspective/1-csvD1MbfQ39AIFmP4Tp1OQ.webp) **Figure 2:** Boltz-2 across four `2xlarge` GPU tiers (median over 5 cases, 8 vCPU held fixed so the GPU is the main variable). Bars show model-compute time and billed time; as the GPU gets faster, model compute shrinks steeply (T4→L40S 4.9×) while billed time shrinks less (2.3×) because a near-fixed container/data overhead does not scale with the GPU. Mean GPU utilization drops accordingly; VRAM stays flat at ≈10 GB, so the workload is not memory-bound. | Instance | GPU | Model (s) | Billed (s) | VRAM (GB) | GPU util (%) | Speedup vs T4 | | ------------ | ---- | --------: | ---------: | --------: | -----------: | ------------: | | g4dn.2xlarge | T4 | 279 | 527 | 9.9 | 81 | 1.0x | | g5.2xlarge | A10G | 88 | 286 | 9.9 | 37 | 3.2x | | g6.2xlarge | L4 | 93 | 276 | 10.1 | 44 | 3.0x | | g6e.2xlarge | L40S | 57 | 230 | 10.5 | 25 | 4.9x | **Table 3: **Boltz-2 across four `2xlarge` GPU tiers (median over 5 TernaryDB cases, 8 vCPU held fixed). Model-compute time is the model process itself; billed time adds the container + data overhead. Results are indicative; a larger case set may shift these numbers. On model-compute time the T4→L40S speedup is 4.9× (279 s → 57 s). On billed time it is only 2.3× (527 s → 230 s): the GPU-independent container/data overhead does not scale with the GPU, so it dilutes the apparent benefit. Isolating model compute shows the GPU helps substantially more than the billed number implies. This is reflected in the utilization numbers. As the GPU gets faster, the inference window shrinks and mean utilization drops: from 81% on the T4 to 25% on the L40S. The workload benefits from faster hardware without requiring larger batches or more memory. **The cost tradeoff.** The T4 (\`g4dn.2xlarge\`, $0.75/hr) has the cheapest hourly rate but, being 3–5× slower, is the MOST expensive per prediction. Multiplying median model-compute time by hourly price, the L4 (\`g6.2xlarge\`) is the lowest cost per prediction at ≈$0.025 — it runs nearly as fast as the L40S at less than half the hourly rate (A10G ≈$0.030, L40S ≈$0.036, T4 ≈$0.058). For batch workloads the L4 is the sweet spot; the L40S buys the lowest latency at a modest premium. **Preprocessing-heavy models are the exception.** For these, a newer instance family can still help, just for a different reason. Running OpenFold3 on the newer \`g6.2xlarge\` rather than the \`g5.2xlarge\` cut median runtime by 1.3× (944 s to 746 s) with VRAM essentially unchanged (≈7 GB), tracking faster CPU-side stages rather than the GPU: median CPU utilization fell from 62% to 52%. Boltz-2 saw only 1.04× from the same swap (Table 3). So for preprocessing-heavy workloads, faster CPU and memory can matter more than a bigger GPU. ## Input Size as a Runtime Proxy For production planning, runtime predictability matters almost as much as raw speed: if you can estimate a job’s runtime up front, you can provision and budget for a batch before launching it. Can you predict how long a job will take before you submit it? For the inference-dominant models in this benchmark, total amino acid count was a strong runtime proxy. The figure below shows model-compute time plotted against total amino acid count (the sum of residues across all protein chains in the complex). ![Model-compute time vs.](https://lokahq.github.io/tech-blog/blog/running-six-open-source-cofolding-models-on-aws-lessons-learned-from-a-compute-perspective/1-V4EHwLYmmNErdTWJBvvqLg.webp) **Figure 3:** Model-compute time vs. total amino acid count for five cofolding models across 29 PROTAC ternary complexes on `g6.2xlarge`. Dashed lines show linear fits for models with r ≥ 0.70. The inference-dominant models track sequence length tightly (Boltz-2, Protenix, RoseTTAFold3, Chai-1); OpenFold3 (r = 0.62) correlates only moderately, with more scatter, consistent with CPU-side variance beyond sequence length. ESMFold2 is omitted: at five samples its cases span two GPU tiers (memory limits), so a single-tier size–runtime fit is not meaningful for it. Total amino acid count is a strong predictor of model-compute time for the inference-dominant models on this dataset: Pearson r = 0.97 for Boltz-2 and Protenix, r = 0.98 for RoseTTAFold3, and r = 0.89 for Chai-1. A small calibration run (a handful of representative cases) may be sufficient to fit this relationship before scaling to larger batches. OpenFold3 shows a moderate correlation (r = 0.62), but with considerably more scatter: in our runs, CPU-side preprocessing cost appeared to vary with factors beyond sequence length alone (likely MSA depth and feature construction load), adding variance that a linear fit on sequence length cannot fully capture. **Practical implication.** For large-scale screening campaigns, a pre-submission cost estimate based on FASTA length can help prevent over-provisioning. This appears reliable for Boltz-2, Protenix, and RoseTTAFold3 on inputs similar to the ones tested here. For Chai-1 and OpenFold3, empirical profiling on a representative sample is a safer approach. ## Cost Per Prediction The headline cost numbers look low, but they should be interpreted carefully. Table 4 combines median model-compute time with public _raw EC2_ on-demand pricing for the \`g6.2xlarge\` instance ($0.978/hr = $0.000272/s) to give an estimated cost per prediction. Note that the benchmark actually ran on SageMaker Processing Jobs (\`ml.g6.2xlarge\`), which are billed at about a 25% premium over the equivalent raw EC2 rate ($1.222/hr on SageMaker vs $0.978/hr raw EC2); we use the EC2 price as a transparent, widely quotable lower bound, so the true SageMaker cost is higher. These are lower-bound estimates in other respects too: they exclude MSA generation, container provisioning (≈49 s, unbilled), the per-run image-pull and data-download overhead (≈150–235 s, billed but amortizable by batching; see below), failed jobs, storage, orchestration overhead, and any idle or queueing time. | Model | Struct./run | Model (s) | Billed (s) | VRAM (GB) | RAM (GB) | Est. cost | Profile | | ------------ | ----------: | --------: | ---------: | --------: | -------: | --------: | ------------------- | | RoseTTAFold3 | 6 | 74 | 311 | 6.9 | 8 | $0.020 | Inference-dominant | | Protenix | 5 | 118 | 300 | 7.3 | 5.4 | $0.032 | Inference-dominant | | Boltz-2 | 5 | 127 | 305 | 10.1 | 9.6 | $0.034 | Inference-dominant | | ESMFold2 \* | 5 | 222 | 356 | 20 | 18.1 | $0.060 | Preprocessing-heavy | | Chai-1 | 5 | 424 | 596 | 12.1 | 12.5 | $0.115 | Inference-dominant | | OpenFold3 | 5 | 493 | 722 | 11.7 | 26 | $0.134 | Preprocessing-heavy | **Table 4: **Per-prediction resource and cost summary on `g6.2xlarge` (L4, 24 GB), median over 29 PROTAC complexes at 5 diffusion samples per run. _Model_ is the model process; _Billed_ adds the container + data overhead SageMaker charges for. _Est. cost_ is model-compute time × raw EC2 on-demand price ($0.978/hr; SageMaker is \~25% higher). _Struct./run_ is the number of structures each model emits, so the cost is per _run_, not per structure. Sorted by model-compute time. Under model-compute accounting, every successful prediction cost under $0.15 on \`g6.2xlarge\` (under $0.20 even counting the full billed overhead). At 10,000 predictions the most efficient models cost roughly $200 (RoseTTAFold3) to $340 (Boltz-2) in model compute, but $830–850 billed at the default per-run overhead, which is exactly why batching to cut that overhead matters at scale, before MSA generation, orchestration, storage, retries, and other overhead. Real production costs will be higher. ![Where the wall-clock goes, per model, on g6.2xlarge .](https://lokahq.github.io/tech-blog/blog/running-six-open-source-cofolding-models-on-aws-lessons-learned-from-a-compute-perspective/1-TLMdZrEKyZukSHrIiWyN3g.webp) **Figure 4:** Where the wall-clock goes, per model, on `g6.2xlarge`. Billed time is the model compute plus a near-fixed container-image pull and input download; provisioning is unbilled. For the fastest models most of the bill is plumbing, not compute, so it is amortizable by batching. For the fastest models, most of the bill is not compute. RoseTTAFold3 spends about 74 s in the model but is billed for roughly 310 s; the remaining \~75% is a near-fixed container-image pull and input download, paid once per instance because we launch a fresh instance per structure (see the timing-breakdown figure above). Batching many structures onto one warm instance amortizes that overhead across the batch and is the single largest cost lever here, larger than the choice of GPU tier. A note on Chai-1: with both models now emitting five diffusion samples, its ≈3.4× longer model-compute time than Boltz-2 (Table 4) is an architectural difference rather than a sampling artifact, and it carries straight into a proportionally higher cost per run. Whether that runtime is justified depends on the accuracy tradeoff, covered in the companion post. ## Conclusions and Deployment Lessons Across 29 PROTAC ternary complexes, the six models split into two profiles: inference-dominant models, whose runtime scaled predictably with input size, and preprocessing-heavy pipelines (ESMFold2 and OpenFold3), where CPU and system RAM, not GPU inference, drove wall-clock time. These are empirical observations from our AWS runs, not a literature-established taxonomy. The practical upshot is that VRAM alone is a poor guide to what an instance needs: OpenFold3 never approached its VRAM ceiling yet needed over 26 GB of system RAM, and a low mean GPU utilization often reflected a short inference burst after a long CPU phase rather than an idle GPU. A few caveats bound all this. It is a compute study only (no biological-accuracy claims, which the companion analysis covers) and limited in scope: a single workload, 29 structures (5 for the instance comparisons), one cloud provider, one seed, a single run per cell with no replicates (so small differences, such as the 1.04× g5→g6, are within cloud noise), and published defaults except a normalised five diffusion samples per model. One model, ESMFold2, could not hold five samples within the 24 GB L4 for most cases and was partly run on a 48 GB L40S. Reported costs are model-computed lower bounds, excluding MSA generation, startup, retries, storage, the SageMaker premium over raw EC2, and idle time, so read them as relative comparisons rather than budgets. With that scope in mind, what we would tell anyone about to run these models at scale: **Don’t pick instances on VRAM alone.** VRAM decides whether a model **can** run, not whether it runs _efficiently_. Match the instance to the workload’s profile: inference-dominant models reward a faster GPU, while preprocessing-heavy ones often gain more from faster CPU and memory. **Runtime is a property of the whole pipeline.** MSA handling, feature construction, input parsing, and container startup all add up, especially for inputs as complex as PROTAC ternary complexes. Profiling only the model, or only the GPU, can misrepresent end-to-end cost. Conversely, billing the whole pipeline can overstate the per-prediction cost of a fast model: for RoseTTAFold3 and Boltz-2 most of the bill is a fixed container and data overhead that a production deployment amortizes by batching, not model compute. We therefore report model-compute time as the primary metric and treat the overhead as a separately-optimizable line item. **Calibrate before you scale.** A small run on representative cases reveals whether runtime tracks input size, whether GPU use is sustained or bursty, and whether CPU or RAM will bind, enough to size instances before launching thousands of jobs. **Compute efficiency is not biological accuracy.** Faster or cheaper is not automatically better. Pair the profiling here with task-specific accuracy validation (in the companion analysis) before choosing a model. **Batch, don’t launch one job per structure.** Because we ran a fresh instance per prediction, a near-fixed container-image pull and data download was billed every time — for the fastest models, that overhead is most of the bill. Batching many structures onto one warm instance amortizes it across the run and is the largest cost lever we found, bigger than the choice of GPU tier. Used this way, these patterns help you size the right instance **before** you scale, rather than discovering the bottleneck after the bill arrives.