PlantAssemblyForge is a modular bioinformatics workflow for plant de novo genome assembly, transcriptome assembly, quality control, assembly validation, functional analysis, and interactive results visualization.
The platform combines Nextflow DSL2 workflow orchestration with established bioinformatics tools and an interactive Streamlit dashboard, providing a reproducible workflow from raw sequencing reads to interpretable assembly-quality results.
PlantAssemblyForge includes an interactive Streamlit dashboard for exploring benchmark results, assembly statistics, validation metrics, k-mer distributions, transcriptome outputs, and workflow architecture.
The overview provides a compact summary of the benchmark dataset and major genome assembly results, including cleaned sequencing coverage, assembly size, genome fraction, and BUSCO completeness.
The genome workflow processes paired-end sequencing reads through quality assessment, read preprocessing, k-mer analysis, assembly, and independent assembly validation.
Paired-end FASTQ
β
βΌ
FastQC
β
βΌ
fastp
β
βΌ
Post-trimming FastQC
β
ββββββββββββββΊ Jellyfish k-mer analysis
β
βΌ
SPAdes
β
ββββββββββββββΊ QUAST
β
ββββββββββββββΊ BUSCO
Representative benchmark statistics include:
| Metric | Result |
|---|---|
| Scaffold assembly size | 107.24 Mb |
| Contig assembly size | 107.21 Mb |
| Scaffold N50 | 8,515 bp |
| Contig N50 | 7,429 bp |
| Largest scaffold | 56,462 bp |
| GC content | 36.21% |
QUAST is used to evaluate assembly contiguity and reference-based genome recovery.
Representative results include:
| Metric | Result |
|---|---|
| Genome fraction | 81.892% |
| Duplication ratio | 1.003 |
| Scaffold NGA50 | 5,971 bp |
QUAST outputs are retained in results_summary/genome/ to provide compact evidence of assembly performance without storing large intermediate files in Git.
BUSCO evaluates recovery of evolutionarily conserved plant genes and provides an independent measure of biological completeness.
Representative BUSCO results:
| Category | Result |
|---|---|
| Complete BUSCOs | 95.4% |
| Single-copy | 93.9% |
| Fragmented | 3.4% |
| Missing | 1.3% |
The benchmark demonstrates high recovery of conserved gene content despite fragmentation expected from a moderate-depth short-read de novo assembly.
Jellyfish is used to generate the k-mer frequency spectrum before assembly.
The current benchmark uses:
k = 21
The resulting spectrum provides information about read multiplicity, sequencing error, coverage structure, and genome complexity.
The representative histogram is available at:
results_summary/genome/k21_histogram.tsv
PlantAssemblyForge also contains modules for de novo and reference-guided transcriptome analysis.
RNA-seq
β
βΌ
FastQC
β
βΌ
fastp
β
βΌ
RNA-SPAdes
β
βΌ
Read-back mapping
β
βΌ
rnaQUAST
β
βΌ
TransDecoder
β
βΌ
DIAMOND
This branch supports transcript reconstruction, assembly validation, protein prediction, and sequence-similarity-based functional analysis.
RNA-seq
β
βΌ
HISAT2
β
βΌ
SAMtools
β
βΌ
StringTie
The reference-guided branch provides an alternative workflow when an appropriate reference genome is available.
PlantAssemblyForge uses modular Nextflow DSL2 processes to separate individual bioinformatics operations from higher-level workflow logic.
Genome workflow modules include:
GENOME_FASTQC
β
βΌ
GENOME_FASTP
β
βΌ
GENOME_FASTQC_CLEAN
β
ββββββββββββββΊ GENOME_JELLYFISH
β
βΌ
GENOME_SPADES
β
ββββββββββββββΊ GENOME_QUAST
β
ββββββββββββββΊ GENOME_BUSCO
The genome modules have also been tested using Nextflow stub execution to validate workflow connectivity without rerunning computationally expensive analyses.
| Category | Tools |
|---|---|
| Workflow management | Nextflow DSL2 |
| Read QC | FastQC, fastp, MultiQC |
| k-mer analysis | Jellyfish |
| Genome assembly | SPAdes |
| Genome validation | QUAST, BUSCO |
| Transcriptome assembly | RNA-SPAdes |
| Alignment | HISAT2, Bowtie2, minimap2 |
| Alignment processing | SAMtools |
| Transcript reconstruction | StringTie |
| Protein prediction | TransDecoder |
| Functional similarity search | DIAMOND |
| Dashboard | Streamlit |
| Data analysis | Python, pandas |
| Visualization | Plotly, Matplotlib |
| Environment management | Conda |
PlantAssemblyForge/
β
βββ app/
β βββ app.py
β
βββ conf/
β βββ genome.config
β
βββ config/
β βββ samplesheet.csv
β
βββ docs/
β βββ images/
β βββ overview.png
β βββ genome_assembly.png
β βββ busco.png
β βββ transcriptome.png
β
βββ modules/
β βββ alignment/
β βββ annotation/
β βββ assembly/
β βββ genome/
β βββ qc/
β βββ validation/
β
βββ workflows/
β βββ genome.nf
β βββ transcriptome.nf
β
βββ results_summary/
β βββ RESULTS.md
β βββ genome/
β
βββ genome_main.nf
βββ main.nf
βββ nextflow.config
βββ environment.yml
βββ CITATION.cff
βββ LICENSE
βββ README.md
Clone the repository:
git clone https://github.com/Pratik-2002-ux/PlantAssemblyForge.git
cd PlantAssemblyForgeCreate the Conda environment:
conda env create -f environment.yml
conda activate plantassemblyValidate the workflow structure using Nextflow stub execution:
nextflow -C conf/genome.config run genome_main.nf -stub-runRun the complete genome workflow:
nextflow -C conf/genome.config run genome_main.nfLaunch the interactive results dashboard:
streamlit run app/app.pyStreamlit will provide a local browser address, typically:
http://localhost:8501
The dashboard provides dedicated views for:
- Project overview
- Genome assembly
- QUAST
- BUSCO
- k-mer analysis
- Transcriptome analysis
- Workflow architecture
Compact representative results are provided under:
results_summary/
Large files are deliberately excluded from Git version control, including:
- Raw FASTQ/SRA sequencing data
- Trimmed FASTQ files
- Large genome assemblies
- Reference genomes and annotations
- BUSCO databases
- Large intermediate alignment files
- Nextflow
work/directories - Tool installations and databases
This keeps the repository lightweight while preserving the workflow implementation and representative evidence required to understand and reproduce the analysis.
PlantAssemblyForge supports reproducibility through:
- Nextflow DSL2 workflow orchestration
- Modular process definitions
- Conda environment specification
- Explicit configuration files
- Representative benchmark outputs
- Git version control
- Versioned releases
- Citation metadata
- Stub-run workflow validation
PlantAssemblyForge v1.1.0
This release includes the modular workflow implementation, representative genome/transcriptome results, interactive Streamlit dashboard, and dashboard documentation.
Citation metadata is provided through:
CITATION.cff
If PlantAssemblyForge contributes to research or analysis, please cite the repository and corresponding software release.
PlantAssemblyForge is distributed under the MIT License.
See LICENSE for details.
Pratik Ramchandra Chaudhari
PlantAssemblyForge was developed as a bioinformatics workflow-development project focused on reproducible plant genome and transcriptome analysis.
β If you find PlantAssemblyForge useful, consider starring the repository.



