Contributing to Open-Source Bioinformatics: Our Experience at the nf-core Hackathon from Medellín

Contributing to Open-Source Bioinformatics: Our Experience at the nf-core Hackathon from Medellín

Written by Andres Florian and Daniel Sabogal

What is nf-core, and why does it matter?#

If you work anywhere near data-heavy science — genomics, metagenomics, proteomics — you’ve probably hit the same wall: building reliable, reproducible data pipelines is hard. Nextflow is an open-source workflow engine designed to solve exactly that problem, letting scientists write scalable pipelines that run consistently across laptops, clusters, and cloud environments.

nf-core is the community layer on top of Nextflow. It’s a collaborative project that curates high-quality, peer-reviewed pipelines and reusable modules — think of them as vetted building blocks that any researcher can plug into their own workflows. The result: less time reinventing the wheel, more time doing actual science.

Twice a year, nf-core organizes a global hackathon where contributors from around the world gather — online and at local hubs — to build, improve, and ship these community resources. In March 2026, the hackathon brought together 29 hubs across 20 countries. Loka hosted two of those hubs: one in Medellín, Colombia, and another in Skopje, North Macedonia.

Why Loka at an nf-core hackathon?#

Nextflow comes up regularly in our client work — in cloud migrations, proof-of-concept builds, and production-grade pipeline development — so contributing back to the open-source ecosystem felt like a natural step. This post covers the Medellín experience.

The Medellín node: small team, focused mission#

Our hub operated out of a WeWork space in Medellín on March 11–12. Three of us worked a standard 9-to-5 with a clear goal: create and publish new nf-core modules.

We scoped five modules and made strong progress on two, taking both to a publishable state within the two days, both of which were later merged into the official nf-core repository.

Module 1: stacks/refmap — variant calling for ddRAD-seq data#

Stacks is a widely used software suite for analyzing restriction site–associated DNA sequencing (RAD-seq) data, a technique commonly used in conservation biology, population genetics, and ecological research. The refmap component performs reference-based variant calling from this type of sequencing data.

During the hackathon we:

  • Refined the module’s process definition to meet nf-core standards.
  • Tested and validated Docker and Singularity containers, as well as the Conda environment definition.
  • Prepared the pull request against the official nf-core/modules repository.

The module was essentially finished by end-of-day two. After the hackathon, a maintainer left some review comments; we addressed them, and the module was successfully merged into the official nf-core repository. You can find it live at nf-co.re/modules/stacks_refmap.

Module 2: singlem/dbdownload— taxonomy database for metagenomics#

SingleM is a tool for metagenomics taxonomy classification — it helps researchers figure out which organisms are present in a complex environmental sample. The module we worked on wraps its database-download functionality, allowing users to fetch the required taxonomic database directly within an nf-core pipeline.

During the hackathon we:

  • Refined the process definition and environment configurations.
  • Fixed linting issues and completed all required nf-test tests.

We finished the pull request shortly after the event, and this module was also merged into the official nf-core repository. You can find it live at nf-co.re/modules/singlem_dbdownload.

What we learned#

Beyond the code itself, the hackathon was a crash course in the nf-core ecosystem. Some highlights:

  • nf-core module standards and best practices. The project has thorough guidelines for everything from naming conventions to container specifications. Working within these constraints taught us a lot about building robust, community-ready bioinformatics resources.
  • Testing with nf-test. Writing meaningful tests for bioinformatics modules — where inputs and outputs are often large, domain-specific files — requires thoughtful test-data selection and snapshot strategies.
  • The power of daily stand-ups. Each day ended with an Americas-region wrap-up session (Day 1 · Day 2) where every hub shared progress. Seeing teams from Buenos Aires, Cambridge, and North Carolina working in parallel was genuinely motivating.

What’s next#

We ran in parallel with Loka’s other hub in Skopje, a team of 20+ tackling larger-scale pipeline work: extending the existing nf-core/proteinfold pipeline with new tools for evaluating the accuracy of AI-based protein structure predictions, and building a brand-new antibody optimization pipeline. Two hubs, two very different scopes, same goal of giving back to the ecosystem.

On our side, two merged modules is a strong start for a first outing, and we’re not done. We still have the rest of our scoped backlog and are continuing to work through it, with more PRs to nf-core/modules planned in the coming weeks.

If you work with data-heavy science and have never contributed to open source, a hackathon is a low-risk way in: the standards are well documented, maintainers are responsive, and you walk away with merged work that the whole community uses.

The Medellín node of the nf-core March 2026 hackathon — a small but determined team
The Medellín node of the nf-core March 2026 hackathon — a small but determined team

If you’re interested in contributing, keep an eye on nf-co.re/events for the next hackathon — they run twice a year.

Originally published on Loka Engineering on Medium.

Tags

Topics