
As next-generation sequencing (NGS) becomes more commonplace in forensic genetics, one of the critical questions facing NGS-equipped laboratories is: how can scientists consistently interpret, share and compare short tandem repeats (STR) sequencing data?
At the center of this challenge is STR nomenclature – the system used to describe STR alleles. Insights from experts such as Dr. David Ballard and Dr. Lucinda Davenport (King’s College London, KCL) and Dr. Kris Van Der Gaag and Jerry Hoogenboom (Netherlands Forensic Institute, NFI) highlight why standardization is not just helpful, but essential.
From length to sequence: Why we need to adapt how we describe STR data
For decades, forensic DNA profiling relied on length-based STR analysis using capillary electrophoresis (CE). In this approach, the amplified fragment is used to infer the structure of the allele, and encompasses both the repeat region and flanking region of the allele, meaning that variation outside the STR itself can influence overall fragment length. Sequence-level variation between alleles of the same size cannot be captured, meaning this information is lost.
By switching to sequencing, laboratories can effectively analyze the desired DNA sequences, improving precision and increasing discriminatory power. However, this added resolution also introduces complexity. Raw sequence strings from NGS are less straightforward to interpret visually, and limitations on the number of characters accepted by bioinformatic pipelines and tools can make it difficult to compare full sequence strings across systems.
Without a standardized system, identical sequences may be described differently by different laboratories or platforms. This creates unnecessary ambiguity and can limit the reliability of comparisons. As the researchers at the NFI explain, a consistent nomenclature is also key to an analyst being able to differentiate between genuine alleles and technical artefacts.
Redefining STRs in a sequencing context
Transitioning to sequence-based analysis is not just a technical upgrade. It requires rethinking how STRs themselves are defined.
One of the main challenges lies in determining where an STR begins and ends within a sequence.
Historical CE definitions are sometimes outdated and do not reflect the underlying sequence, which can lead to inconsistencies.
A shared goal: One language for STR sequencing
The Nomenclature Project of the DNA Commission of the International Society for Forensic Genetics (ISFG) aims to bring consistency to how sequence-based STR data are reported. While the objective is straightforward, the impact is far-reaching. This initiative will allow laboratories to:
- Compare results across workflows and platforms
- Distinguish genuine alleles from technical artifacts
- Support compatibility with existing DNA databases
This shared language supports collaboration, enables smoother data exchange and reduces the risk of misinterpretation.
Turning sequence data into usable information
Even among experts, agreement on loci naming is not always easy. The NFI team notes that different opinions can arise when defining optimal allele names, even within the same group. Some respondents often leaned to longer names that included a larger part of sequence of complex STR loci.
This is why community-driven consensus is so important. Clear, agreed guidelines ensure that naming conventions remain consistent and practical across laboratories.
Another challenge is making sequence data usable for databases. Raw sequences must be translated into standardized, human-readable allele names for routine comparison.
As researchers from KCL emphasize, this conversion is essential for comparing reference and casework samples. Without standardization at this stage, the value of sequencing data is significantly reduced.
By bridging sequence data and database formats, standardized nomenclature ensures compatibility with existing systems while supporting future developments.
Supporting clearer mixture interpretation
Mixture interpretation remains one of the most demanding aspects of forensic analysis. Sequence-based data provides more detail, and can significantly improve deconvolution of profiles thanks to the additional variants observed, but without structure, that detail can be difficult to interpret.
Standardized nomenclature helps bring clarity. By clearly representing sequence variation, it becomes easier to distinguish between stutter artefacts and true minor contributors. This leads to a more confident interpretation and strengthens the reliability of results.
Why standardization makes a difference
Standardized nomenclature changes how laboratories work with sequencing data. Instead of creating their own naming conventions, scientists can rely on an established framework that supports consistency from the start.
This makes it easier to compare results across laboratories, even when different kits or workflows are used. It also simplifies the adoption of NGS, as laboratories can integrate sequencing without building entirely new interpretation systems.
At the same time, the technical burden is reduced. Laboratories can shift their focus from defining systems to applying and verifying them, saving time and improving efficiency.
Making standardization part of the workflow
For nomenclature standards to have a real impact, they must be integrated into everyday tools and processes.
This includes software that applies standardized naming automatically, outputs that follow consistent formats and clear documentation of how results are generated as well as a way of referring back to the original sequence when needed for transparency. As the NFI team points out, it is critical to avoid situations where laboratories unknowingly compare data generated using different naming systems.
Encouragingly, solutions such as FDSTools and Universal Analysis Software for ForenSeq Signature Plus are already supporting these standards, helping laboratories apply them consistently in routine work.
Where we are today
The harmonized STR nomenclature is being implemented among laboratories that sequence STRs. Key journals in forensic genetics require STR sequence data to be submitted in a format that follows the recommendations of the DNA Commission of the ISFG, and software solutions are adapting their input requirements to the standardised allele names.
By combining high-resolution sequencing with consistent naming rules and appropriate validation, laboratories can adopt NGS technologies while preserving long-established data systems.
Standardization is becoming the foundation for reliable data sharing, consistent interpretation and the continued evolution of forensic databases.
A shared language for the future of forensics
Consistent STR nomenclature turns complex sequence data into something laboratories can read, compare and act on with confidence.
Without standardization, the full value of sequencing cannot be realized. With it, forensic genomics becomes more connected, more consistent and better equipped to support investigations worldwide.
People we interviewed about the nomenclature project

Forensic Scientist, Netherlands Forensic Institute



Author
