Contents
- Introduction: Beyond Sequencing
- Workflow Design vs Real-World Constraints
- Sample Logistics: A Layer with Specific Fragilities
- Sample Identification and Naming Systems in Newborn Genomic Screening
- Data Infrastructure and Workflow Management
- Turnaround Time in High-Throughput Clinical Settings
- From Screening to Clinical Follow-Up
- Conclusion: Newborn Genomics as an Operational System
This document outlines key operational challenges encountered in large-scale genomic screening programs, with a focus on real-world implementation constraints.
1. Introduction: Beyond Sequencing
In large-scale genomic screening programs, sequencing is often perceived as the central technical challenge. In practice, however, the most critical phase occurs before sequencing even begins.
The successful implementation of genomic newborn screening requires the alignment of a complex network of stakeholders, each with distinct roles, expectations, and levels of influence. These typically include:
- hospital clinical leadership
- research governance and ethics committees
- bioinformatic and analytical teams
- external genomic centers
- funding bodies and institutional partners
- regulatory and institutional stakeholders
In most cases, genomic newborn screening programs operate within a research setting, where approval by the ethics committee enables the use of early neonatal WES and WGS within clearly defined operational boundaries. This allows the extension of genomic analysis to asymptomatic newborns, which in a purely diagnostic setting would typically be restricted to cases with a defined clinical indication.
This distinction has direct implications for workflow design, reporting structures, and regulatory exposure.
Large-scale newborn screening programs are shaped not only by technical factors, but also by non-technical constraints. In practice, this includes institutional positioning, territorial ownership of the project, and alignment with broader healthcare and political priorities, which can influence how different components of the program are presented and attributed across institutions.
For this reason, genomic screening programs cannot be reduced to sequencing pipelines. Their execution depends on coordinating technical processes with organizational structures, regulatory constraints, and decision-making layers that directly affect how the workflow operates in practice.
From an implementation perspective, this leads to a key principle:
Expertise in sequencing is necessary, but not sufficient.
Execution depends on navigating complex stakeholder ecosystems while maintaining control over critical components of the workflow.
In early-stage programs, strong alignment between stakeholders, often driven by initial enthusiasm, can temporarily mask structural tensions. However, long-term scalability requires a more robust balance between collaboration, autonomy, and operational control.
In the context of genomic newborn screening, different technical approaches may be adopted depending on the design of the program and its clinical objectives. These typically include:
- Whole Exome Sequencing (WES)
- Whole Genome Sequencing (WGS)
- WES-based panels targeting predefined sets of clinically actionable genes
Each approach involves trade-offs in terms of data volume, interpretative complexity, turnaround time, and clinical scope.
WES-based panels, in particular, represent a pragmatic compromise in early-stage or large-scale programs, focusing on genes with established clinical relevance while maintaining manageable data complexity and reporting timelines.
Regardless of the chosen approach, the operational challenges associated with large-scale implementation remain substantial and extend far beyond sequencing itself.
Newborn genomic screening is not a laboratory process. It is a coordinated effort across institutions, roles, and responsibilities.
2. Workflow Design vs Real-World Constraints
When genomic screening programs reach volumes of hundreds of samples per month, workflow design is no longer a secondary consideration. It becomes a central operational problem.
At this scale, execution requires the combination of operational capacity, technical expertise, process precision, and coordination across multiple systems. None of these elements, taken alone, is sufficient.
In genomic newborn screening, timing is a defining constraint. Delivering results within a clinically meaningful window is essential to preserve the value of early detection. In practice, this means returning results within approximately four weeks. Beyond this threshold, the distinction between newborn and postnatal screening becomes clinically and operationally relevant.
For a subset of genetic conditions, even shorter turnaround times would be desirable—such as 7–10 days—to enable timely therapeutic intervention. While this is not yet easily achievable in large-scale screening programs, it represents a key target for the next generation of newborn genomic screening, requiring workflows to be designed with this level of performance in mind.
At this scale, the main constraint is not sequencing itself, but the ability to organize the entire system around delivery timelines.
From an operational perspective, this leads to a central requirement:
Workflow design must be driven by delivery timelines, not sequencing capacity alone.
In practice, this means building systems that are simple to operate, even when the underlying processes are complex.
A key component is the data infrastructure. The practical design of a large-scale program does not start from sequencing, but from the construction of a structured database capable of managing:
- sample identifiers
- sequencing orders
- clinically relevant metadata (e.g. birth weight, head circumference, consanguinity)
Importantly, clinical data is not always available at the time of sample collection. Systems must therefore support asynchronous data integration, allowing information to be added or updated after sample intake, during analysis, or even after reporting, particularly in the context of study-level statistical analysis.
For this reason, simple tabular tools are not sufficient. A robust database architecture is required, capable of handling large volumes of data with consistency and traceability over time.
Real-world constraints further shape the workflow.
These include:
- sequencing capacity and the need for operational redundancy
- sample transport logistics
- high-volume data transfer (e.g. ~15–20 GB per sample in WES)
- secure and scalable data storage, often requiring cloud-based solutions
- integration with bioinformatic analysis platforms
Downstream analysis introduces an additional layer of complexity. At scale, the bottleneck shifts from sequencing to interpretation. Bioinformatic pipelines must be designed to process large volumes of data without amplifying the interpretative load.
Clinical interpretation then becomes a point of responsibility. Results are translated into structured reports, validated and signed by a medical geneticist — accountable not only for what is included, but also for what is deliberately left out.
At scale, workflow is no longer defined by sequencing capacity, but by the ability to deliver results within a clinically meaningful timeframe.
3. Sample Logistics: A Layer with Specific Fragilities
In high-throughput genomic screening programs, sample logistics is often underestimated during the design phase. Standardized formats such as 96-well plates are commonly perceived as efficient and scalable solutions, particularly when aligned with sequencing platform requirements.
In controlled environments, this assumption may hold true. In real-world multi-center projects, however, logistical constraints introduce variables that are rarely accounted for during initial planning.
A typical requirement is the organization of incoming samples into 96-well plates prior to sequencing. This implies not only precise sample tracking and labeling, but also physical stability during transport across different facilities and conditions.
In practice, this approach introduces critical risks.
Transport conditions—temperature fluctuations, mechanical stress, and handling variability—can compromise plate integrity. Sealing films, while sufficient for in-lab workflows, may not withstand shipping conditions. The result can be cross-contamination between wells, leading to compromised data quality and the need to repeat entire batches.
The solution was deceptively simple: abandoning plate-based transport in favor of individually sealed sample tubes, grouped and tracked at the system level rather than physically constrained into fixed plate formats.
This shift reduced contamination risk, simplified handling, and improved robustness across the entire pipeline—without affecting downstream sequencing processes.
The key insight is that workflow design must prioritize real-world robustness over theoretical efficiency. What appears optimal in a controlled laboratory setting may fail when exposed to variability in transport, handling, and coordination across multiple actors.
Scalability in genomic screening is not achieved by forcing samples into predefined structures, but by designing systems that remain stable under real-world conditions.
4. Sample Identification and Naming Systems in Newborn Genomic Screening
In large-scale newborn genomic screening programs—whether based on WES, WES-based panels, or WGS—sample identification is not a secondary technical detail. It is a foundational component of the entire system.
Most of these programs still operate within a research framework rather than routine clinical practice. As a result, strict pseudonymization rules apply. Patient identity must remain fully separated from genomic data throughout the entire workflow.
In practice, this means that the laboratory performing the genomic analysis does not have access to patient names or direct identifiers. The association between sample ID and patient identity is maintained exclusively within the hospital, often in restricted and non-networked systems.
The laboratory receives and processes only coded samples. Final reconciliation between genetic results and patient identity is performed by the hospital before delivering results to the family.
This separation introduces a critical operational challenge: the reliability of the sample identification system across the entire workflow.
In many settings, sample codes are generated locally within the hospital. However, this approach is inherently fragile. Samples often pass through multiple hands, departments, and intermediate steps before reaching the genomic laboratory. Each transition increases the risk of mislabeling, duplication, or data inconsistency.
For this reason, we believe that sample identification should be generated centrally by a dedicated digital system provided by the laboratory.
In our experience, sample identification should be centrally managed by the genomic laboratory performing the screening. Generating and controlling sample identifiers within the laboratory system ensures full consistency from initial processing to final reporting.
This also enables the use of a single, structured and pseudonymized database, supporting reliable data tracking and downstream statistical analysis.
This approach allows both the laboratory and the clinical partner to track each sample from collection to analysis, reporting, and even downstream statistical analysis, all within a unified system.
Designing the naming system itself is equally critical.
The identifier must be:
- unique
- simple to generate automatically
- easy to read and transcribe
- scalable to the expected volume of the project
A common mistake is to use naïve numeric systems that do not behave correctly when sorted computationally. For example:
1
12
3
In standard alphanumeric sorting, “12” appears before “3”.
To avoid this, identifiers should follow structured formats such as:
SMP0001
SMP0002
SMP0003
This ensures correct ordering across databases, spreadsheets, and downstream analytical systems.
Finally, the physical labeling of samples must reflect real-world constraints. While printed labels are ideal, manual labeling is often unavoidable in high-throughput clinical settings. For this reason, identifiers must be simple and short enough to be easily handwritten, minimizing the risk of ambiguity.
Simple and short identifiers are also critical in downstream wet lab phases. Sequencing platforms often enforce strict input masks, limiting identifier length, formatting, and the use of spaces or complex character patterns. In practice, the simpler the identifier, the more robust the entire process becomes.
If you don’t control the sample ID, you don’t control the system.
5. Data Infrastructure and Workflow Management
Data infrastructure is the operational backbone of any large-scale genomic newborn screening program.
All data must be managed within a proper structured database system designed to ensure consistency, traceability, and long-term integrity. At this scale, tools such as Excel or text-based systems are not just suboptimal—they introduce a concrete risk of data corruption, duplication, and loss of control.
For this reason, genomic workflows must rely on robust database systems such as MySQL, PostgreSQL, or equivalent relational databases, capable of handling complex and high-volume operations.
The responsibility for this infrastructure must reside within the genomic laboratory, which is the only entity that follows the sample throughout the entire process—from sample identification to sequencing data, variant interpretation, and final reporting.
A properly designed system must allow integrated management of:
- sample identifiers
- sequencing orders and processing status
- genomic data and clinically relevant variants
- report generation and storage
An important practical constraint is that clinical data is not always available at the time of sampling. Therefore, the system must support asynchronous data integration, allowing clinical information to be added or updated at later stages—after sample shipment, during analysis, or even after reporting, when final statistical analyses are performed.
Report Generation and Interpretation
The ideal report is generated directly from the bioinformatic workflow.
Bioinformatic platforms should automatically transfer structured data—such as sample identifiers and selected variants—into report templates, ensuring stability and reducing manual errors.
At the same time, reporting must remain semi-automated. Structured where possible, but open to expert input where necessary.
Even in the absence of detailed clinical information, which is common in newborn screening, certain variants require careful interpretation, clear wording, and clinical judgment. The reporting system must allow this flexibility, enabling geneticists to add case-specific comments that are understandable and clinically meaningful for the receiving physician.
A critical layer of maturity is reached when interpretation outputs are systematically archived and made reusable.
In WES-based panels, due to the limited number of genes, the same variants are frequently observed across multiple samples. Storing previously curated interpretations allows for:
- consistent reporting across cases
- reduction of turnaround time (TAT)
- improved operational efficiency
- reduced cognitive load and interpretative fatigue for clinical scientists
This is not just an efficiency gain—it is a key component of methodological consistency.
Batch-Based Workflow Organization
In newborn screening programs, samples are collected and processed in batches, typically reflecting the clinical volume of the originating institution.
In practice, samples are transferred to the genomic laboratory in regular batches, often on a weekly basis depending on clinical volume and organizational setup.
For this reason, database systems must support not only individual sample tracking, but also batch-level identification and management.
Batch identifiers can be used to:
- organize sequencing runs
- track sample flow across the pipeline
- monitor shipment and processing status
- identify failed or incomplete samples
Batch-based organization also simplifies communication with clinical partners.
Workflow Design and Human Factors
High-throughput genomic screening involves tasks that are technically demanding but also repetitive, particularly at the level of variant interpretation.
For this reason, workflow design must take into account not only technical processes, but also human workload and cognitive burden.
Laboratory teams must be supported by:
- clear and structured workflows
- intuitive bioinformatic platforms
- systems that reduce unnecessary repetition
- access to previously curated interpretations
The goal is not to replace expert judgment, but to allow it to be applied more efficiently, more consistently, and with less friction across large volumes of samples.
Result Delivery and Clinical Communication
The final step of the workflow is the delivery of results to the clinical partner.
To avoid inefficient workflows on the clinical side, results should be delivered in a structured, batch-based format, with clear indication of which samples carry clinically relevant findings. This prevents clinicians from having to manually review each individual report to identify positive cases and allows rapid recognition of those requiring immediate attention.
In genomic newborn screening, reporting is typically limited to pathogenic and likely pathogenic variants, in accordance with current clinical guidelines (ACMG). From a clinical perspective, this distinction is largely technical, as both categories are managed in the same way.
Consistency in genomic screening is built in the database, not only in the report.
6. Turnaround Time in High-Throughput Clinical Settings
Turnaround time (TAT) is a defining parameter in genomic newborn screening.
To be consistent with the definition of “newborn screening”, genetic reports must be delivered within a clinically relevant timeframe, typically within 3–4 weeks from birth.
This is not an unrealistic target. It is achievable when the workflow is properly structured, coordinated, and designed with timing in mind from the beginning.
Reports delivered beyond this timeframe may still provide valuable genetic information. However, they no longer fulfill the functional role of newborn screening and instead fall into the domain of perinatal or delayed postnatal diagnostics.
Looking ahead, further reductions in turnaround time will become increasingly important. For a subset of time-critical genetic conditions, earlier results can directly impact clinical management and access to emerging therapies.
In these cases, an ideal target is a TAT of 7–10 days, which represents a meaningful threshold for enabling early intervention.
In high-throughput genomic screening, turnaround time is not just an operational metric. It is a clinical variable that directly influences the value of the test.
Today TAT: 3–4 weeks.
Next: 7–10 days
7. From Screening to Clinical Follow-Up
Once the genomic screening report is delivered, responsibility shifts to the clinical team.
At this stage, interaction between the genomic laboratory and the hospital becomes critical. Even when reports are clearly written, clinicians frequently require clarification on how to act on specific variants, particularly in cases where interpretation is complex or implications are uncertain.
This is where close collaboration between laboratory and clinicians becomes essential, with feedback provided within a short timeframe.
At the same time, this interaction reflects a fundamental asymmetry of responsibility.
The laboratory carries the burden of deciding whether a variant should be reported and how it should be interpreted, often in the absence of complete clinical information.
The clinical team, on the other hand, is responsible for communicating results to families and managing potential overmedicalization, particularly in cases where findings may be uncertain or of variable penetrance.
For this reason, alignment between laboratory and clinicians must be continuous, not episodic.
Longitudinal Value of Genomic Data
Genomic screening does not end with report delivery.
In sequencing-based approaches, particularly WES and WGS, the generated data retains long-term clinical value. Patients may develop symptoms months after birth, or new knowledge may emerge that changes the interpretation of previously identified variants.
For this reason, laboratories must be able to:
- retrieve archived sequencing data
- rapidly reanalyze cases when clinically indicated
- extend analysis beyond the initial scope when necessary
This includes the ability to upgrade analysis, for example, expanding from targeted panels to full exome or genome interpretation when new clinical indications arise.
In genomic medicine, the report is not the end of the process. It is the beginning of clinical management.
8. Conclusions: Newborn Genomics as an Operational System
Newborn genomic screening is often described in terms of sequencing technologies, gene panels, or analytical pipelines.
In practice, however, its success depends on something more fundamental: the ability to design and orchestrate a coherent system.
From sample identification to data infrastructure, from workflow organization to turnaround time, each component contributes to the overall reliability of the process. In complex systems, imperfections are inevitable. What defines success is the strength of the underlying structure and the ability to continuously adapt and resolve them.
For this reason, newborn genomics should not be approached as a collection of technical steps, but as an integrated operational system that must be carefully orchestrated.
Technology enables genomic screening. Coordination makes it work.
Newborn genomics is not a concept.
It is an operational system.
This is how we design it.