# From Proteins to Pipelines: How Open-Source Nextflow Tools Are Accelerating Drug Discovery > Modern computational drug discovery depends on complex workflows that connect AI models, structural biology tools, databases, and quality-control steps into reproducible pipelines. While many organizations have access… - URL: https://lokahq.github.io/tech-blog/from-proteins-to-pipelines-how-open-source-nextflow-tools-are-accelerating-drug-discovery/ - Type: Blog article - Authors: Jelena Pejovic (Senior Bioinformatic Engineer), Jorge Moura Sampaio (Bioengineering Lead) - Published: 2026-07-08 - Updated: 2026-09-30 - Reading time: 6 min - Tags: Nextflow, Machine Learning, Drug Discovery, Bioinformatics, Open Source - Originally published on Medium: https://medium.com/loka-engineering/from-proteins-to-pipelines-how-open-source-nextflow-tools-are-accelerating-drug-discovery-98f25e6993f4 --- _Written by Jelena Pejovic, Senior Bioinformatic Engineer and Jorge Moura Sampaio, Bioengineering Lead_ Modern computational drug discovery depends on complex workflows that connect AI models, structural biology tools, databases, and quality-control steps into reproducible pipelines. While many organizations have access to state-of-the-art models, deploying and scaling them in production remains a significant engineering challenge. At the nf-core Hackathon March 2026 in Skopje, North Macedonia, Loka’s Bioinformatics and Bio ML teams contributed two open-source Nextflow projects that demonstrate how we help organizations operationalize computational drug discovery workflows. One extends protein structure prediction pipelines with standardized evaluation capabilities, while the other provides an end-to-end workflow for computational antibody engineering. Built alongside students from the Faculty of Computer Science and Engineering in Skopje, these projects showcase the type of production-grade workflow infrastructure that enables teams to move from individual tools and models to scalable, reproducible drug discovery platforms. ## Why Nextflow Is the Infrastructure Layer for Modern Drug Discovery If your organization runs computational biology workloads, chances are you’re already using [Nextflow](https://www.nextflow.io/). The numbers tell the story: over 20,000 bioinformaticians use it worldwide, it powers 40% of modern biology workflows, and it runs 2.5 million pipeline executions every month. Sixty-five percent of the world’s top pharmaceutical companies have adopted it. On top of Nextflow sits nf-core — a community-maintained collection of over 1,400 peer-reviewed modules and 100+ full pipelines. When a module lands in nf-core, it’s available to every researcher and every team globally. That’s the model: build once, share with everyone, improve collectively. What makes this ecosystem particularly powerful right now is how naturally it bridges machine learning and biology. The tools predicting protein structures — AlphaFold, Boltz, ColabFold — are deep learning models. The tools evaluating and validating their outputs are software engineering problems. Nextflow is the orchestration layer that connects all of it: running each tool in its own container, parallelizing across thousands of samples, and producing reproducible results on any infrastructure from a laptop to AWS. For organizations building drug discovery capabilities, the gap between a breakthrough ML model and a production-ready workflow is an engineering problem. Drug discovery workflows often involve dozens of tools, data sources, and validation steps that must run reliably across local infrastructure, HPC environments, or the cloud. That’s the problem Loka solves. This is where workflow orchestration becomes critical. Nextflow provides the foundation for building reproducible and scalable computational drug discovery platforms, allowing teams to automate complex processes while maintaining consistency across projects and environments. ## Projects Overview ## Project 1 — Protein Complex Co-folding Protein structure prediction is increasingly becoming a workflow problem rather than a modeling problem. Modern teams routinely generate thousands of structures using tools such as AlphaFold, Boltz, and ColabFold, but extracting actionable insights requires automated evaluation and quality-control steps that can scale alongside prediction. Our contribution extends the nf-core/proteinfold workflow with structural evaluation modules, enabling researchers to automatically assess prediction quality as part of the same reproducible pipeline. ## Project 2 — Computational Antibody Engineering The challenge in computational antibody engineering is not a lack of tools. Researchers already have access to powerful models for sequence design, structure prediction, humanization, and developability assessment. The challenge is connecting these tools into a single reproducible workflow that can be executed consistently across projects and scaled from individual candidates to large design campaigns. Most organizations still rely on fragmented processes involving multiple scripts, environments, and manual handoffs. Our second project addresses this workflow orchestration challenge directly. ### Project 1 — Structural Evaluation for nf-core/proteinfold ### The Pipeline [nf-core/proteinfold](https://nf-co.re/proteinfold/2.0.0) unifies multiple protein structure prediction tools — AlphaFold2, AlphaFold3, ColabFold, ESMFold, Boltz, and others — under a single reproducible workflow. It solves a fragmentation problem: each tool has different dependencies, input formats, infrastructure requirements, and output conventions. Proteinfold standardizes all of that. The pipeline already reports per-model confidence metrics like pLDDT through MultiQC. What it lacked was the ability to compare predicted structures against external reference models, and to score the quality of binding interfaces in protein complexes specifically. That’s what our two modules add. ### What We Built We added two evaluation modules that answer the question every downstream decision depends on: how reliable is this prediction? **US-align — structural comparison.** US-align measures how similar a predicted structure is to a known reference, atom by atom. It produces RMSD (a 3D difference score where lower is better) and TM-score (a global similarity metric from 0 to 1, where closer to 1 is better). For any team using AI-predicted structures to guide experimental work, these metrics are essential for deciding whether a prediction is actionable. **DockQ v2 — interface quality scoring.** When two proteins fold into a complex, the binding interface is what matters most. DockQ scores that interface specifically, producing metrics including interface RMSD, ligand RMSD, fraction of native contacts, and an overall DockQ score. If you’re designing molecules that need to bind a specific surface, this is the metric that tells you whether your prediction got the critical part right. Together, US-align and DockQ complete proteinfold’s evaluation story: predict the structure, compare it to a reference, and score the binding interface. By integrating these capabilities directly into nf-core/proteinfold, structural evaluation becomes a reproducible step within a larger drug discovery workflow. ## Project 2 — An End-to-End Antibody Optimization Pipeline ### The Problem Researchers working on computational antibody engineering typically chain steps manually across disconnected tools and environments. The process is hard to reproduce, difficult to scale, and creates friction between computational design and wet lab validation. Our second project addresses that directly with a single Nextflow pipeline that takes a known antibody structure and produces a ranked list of optimized candidates ready for laboratory testing. ### What We Built ![Nextflow workflow from antibody CDR redesign with AntiFold through structure verification, humanization, and humanness scoring.](https://lokahq.github.io/tech-blog/blog/from-proteins-to-pipelines-how-open-source-nextflow-tools-are-accelerating-drug-discovery/1-Ob4jubvpMXTp1vq8XELTyg.webp) Nextflow workflow for antibody CDR redesign, structure verification, humanization, and humanness scoring. **Step 1 — AntiFold: CDR Redesign via Inverse Folding.** Standard folding goes from sequence to structure. AntiFold does the reverse — it takes the 3D backbone of an antibody’s CDR and generates hundreds of new sequences compatible with that shape. Instead of testing trillions of possible CDR combinations, you constrain the search to sequences that are physically compatible with the binding geometry you need. **Step 2 — ABodyBuilder2: Structural Verification.** Not every sequence AntiFold generates will fold correctly in practice. ABodyBuilder2 takes each candidate and predicts its 3D structure independently — essentially refolding from scratch to verify it still produces a valid antibody. Think of it as a compiler check: you generated the code, now you run it to confirm it compiles. **Step 3 — BioPhi Sapiens: Humanization.** Most research antibodies originate from mice. A mouse antibody in a human patient triggers an immune response. Humanization rewrites the non-CDR portions to resemble human antibodies as closely as possible, while preserving the designed binding site. **Step 4 — OASis: Humanness Scoring.** OASis scores each humanized candidate against a database of over one billion real human antibody sequences. It breaks the sequence into overlapping 9-residue windows and checks how many appear in actual human antibodies. Higher scores mean the sequence looks more human and is likely safer in a patient. By orchestrating the entire process — from design to validation and humanization — the pipeline transforms a fragmented computational process into a reproducible and scalable antibody engineering platform. ## From Models to Production Workflows The projects developed during the hackathon highlight a broader trend across computational drug discovery: the bottleneck is increasingly workflow infrastructure rather than individual algorithms. Modern pipelines combine structure prediction, molecular design, sequence generation, quality control, and downstream analytics. While many organizations have access to powerful AI models, integrating them into reliable, scalable workflows remains a significant challenge. Nextflow addresses this challenge by providing a framework for orchestrating complex scientific workflows across any infrastructure. The projects described here are examples of how Loka helps organizations transform computational methods into reproducible workflows that scientists can run, validate, and scale. Both projects will continue to evolve through the nf-core review process and future benchmarking efforts. More importantly, they demonstrate the type of workflow infrastructure required to operationalize modern computational drug discovery. At Loka, we help organizations build and deploy these workflows using Nextflow and modern cloud infrastructure, enabling teams to move faster from computational insight to experimental validation.