Skip to content

LoReRNA v1.0.0 quality hardening, provenance and results explorer - #3

Merged
dariusng28-collab merged 27 commits into
mainfrom
release/container-v1.1.0
Aug 6, 2026
Merged

dariusng28-collab merged 27 commits into
mainfrom
release/container-v1.1.0

Conversation

@dariusng28-collab

Copy link
Copy Markdown
Owner

Brings main up to the validated release branch. Pipeline stays 1.0.0,
containers 1.1.0.

Validated by a cold from-scratch run.

  • nf-core layout, nf-schema validation, native resourceLimits
  • CI: nf-core lint, -stub-run DAG check, container CLI contract
  • Software versions in the MultiQC report
  • Retry on exit 138; fail fast on unusable Singularity build scratch
  • lorerna-explorer: Shiny app for DGE/DTE/DTU with transcript structure

dariusng28-collab and others added 27 commits July 22, 2026 12:09
Tier-1 hardening (branch only, additive to runtime):
- CI: lint + container CLI contract on every push/PR
- container_cli_contract.sh: catches tool/flag breaks (oarfish-class) pre-run
- stub: blocks on all 13 processes (only used under -stub-run; script: untouched)
- conf/test.config + make_stub_inputs.sh: -stub-run wiring profile (inert)
- untrack conf/ucl.config (institution /SAN paths; stays local as provenucl.config)
Replace the single 13-column MisER bar (mismatched scales) with a
self-describing MultiQC multi-dataset bargraph: one metric per view
(pass rate, read events, unique micro-exons, exon support, median exon
size), samples on the x-axis. IsoQuant, SWISH and all scientific steps
unchanged; merged TSV still published, MULTIQC now reads the JSON.
Semantic per-category colours (green=passed/good, red=failed, blue tiers
for exon support, accents for the single-metric views) to replace the
flat default palette. Safe: MultiQC falls back to defaults if it ignores
pconfig.colors — the report still renders either way.
…colour, IsoQuant-style)

MultiQC ignores custom colours on multi-dataset switcher JSON, so replace it
with one small TSV per metric group + a coloured custom_data block each — the
same mechanism the isoquant_assignment section uses. Pass rate / read events /
micro-exon support / median size, each on its own scale, samples on x-axis.
Add nextflow_schema.json (37 params, typed with enums/defaults) and
assets/schema_input.json; replace ad-hoc checks in main.nf with
validateParameters(). Pin every conda/pip dependency to the versions in the
validated 1.1.0 images so a rebuild reproduces them exactly.
…ates

Tranche 1b of the nf-core-quality work. Pure additions, no runtime impact.
Move process resources to conf/base.config and per-process settings to
conf/modules.config; move workflow LORERNA into workflows/lorerna.nf, leaving
main.nf as a thin entry point. Process names and scripts unchanged, so resumed
runs stay cached.
Move all 13 publishDir declarations out of the module files and into
conf/modules.config, following the nf-core convention of keeping the
published layout in one place.

Preserved verbatim:
  ISOQUANT     saveAs flattening of the nested <sample>/<sample>/ output
  SWISH_PLOTS  saveAs stripping of the pub_plots/ prefix
  CLEAN_BAM    pattern *.clean.bam

MISER container routing merged into its own withName block alongside
its publishDir; no duplicate selector remains.

publishDir is not part of the Nextflow task hash, so this is cache-safe.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add .nf-core.yml declaring this as a pipeline repository, with lint
exclusions for checks that only apply to pipelines hosted under the
nf-core organisation (their CI, branding, AWS megatests, email
templates, RO-Crate). Each exclusion carries its reason inline.
All substantive checks remain enabled.

Real fixes:
  manifest.homePage    was a hard lint failure; also what nextflow info shows
  manifest.mainScript  lint warning
  conf/base.config     add default process cpus/memory/time — previously an
                       unlabelled process would inherit Nextflow's defaults
                       (1 CPU, unlimited memory/time), which SGE mis-sizes

Supporting files: .gitattributes, .prettierrc.yml, .prettierignore,
docs/README.md, .github/ISSUE_TEMPLATE/config.yml.

Remove the redundant .first() on PREPARE_OARFISH_REFERENCE.out.merged_fa —
both process inputs are value channels, so the output already is one.
The .first() on collectFile is required and is unchanged.

Not addressed, pending a decision: params.max_cpus/max_memory/max_time
are deprecated in favour of process.resourceLimits, which would require
raising nextflowVersion to >=24.04.0.
resourceLimits migration (Nextflow >= 24.04):
  Replace the hand-rolled check_max()/params.max_* clamping with native
  process.resourceLimits. Removed check_max() and the max_cpus/max_memory/
  max_time params from nextflow.config, base.config and nextflow_schema.json.
  Each site config (sge/slurm/lsf/test) now sets its own resourceLimits.

  base.config deliberately carries NO resourceLimits — the cap is a
  site-specific fact set by the -c config. A run with no site config is left
  uncapped, exactly as before, so an existing local site config that sets
  neither max_* nor resourceLimits behaves identically.

  Raises the Nextflow floor to >=24.04.0 (HPC runs 25.04.7).

Phase 2 — stub-run CI:
  Add a self-contained 'stub' profile (includeConfig conf/test.config) and
  enable the DAG-wiring job in ci.yml: bash test/make_stub_inputs.sh then
  nextflow run main.nf -profile stub -stub-run. Every module has a touch/mkdir
  stub whose outputs satisfy its output globs, so this catches channel/emit/
  publishDir errors in seconds with no real tools or data.
The publication PDFs are capped by --swish_top_n and split DGE, DTE and
DTU across three files that nothing joins. This adds a Shiny app reading
results/06_swish/results/ that shows one gene across all three analyses
at once, which is the interpretive case the PDFs cannot cover.

Runs locally against the user's own output; nothing is transmitted.
Result CSVs are gitignored by name because the existing results/ rules
are directory patterns and do not match a loose CSV.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The explorer is deployed to shinyapps.io, so 'nothing is transmitted
anywhere' no longer holds. Files are read by whichever machine runs the
app: run locally nothing leaves the machine, run from the hosted copy
they reach a third-party server and are held for the session.

Adds a notice above the file picker in app.R and corrects the wording in
README.md, docs/output.md and lorerna-explorer/README.md. The wording is
accurate for both deployments, since the same app.R serves each.

The hosted URL is not documented yet; it lands once the final app name
is deployed.
The explorer is deployed at https://darius28.shinyapps.io/lorerna-explorer/
and needs nothing installed, which is the lowest-barrier route for
colleagues who do not run R.

Documented in README.md, docs/output.md and lorerna-explorer/README.md,
in each case alongside the existing note that the hosted copy processes
uploads on a third-party server.
Adds MISER_VERSION and SOFTWARE_VERSIONS, which query the tools inside
each of the two images and emit a Software Versions section for the
MultiQC report plus pipeline_info/software_versions.yml.

One dump per image rather than nf-core's per-process versions.yml:
LoReRNA runs every process from one of two images, so per-process files
would repeat the same versions while invalidating the task hash of all
13 analysis modules. MISER in particular does not report its own version
— adding an output there would force every MisER job to re-run.

Both formats are generated from a single pipe-delimited record list, so
adding a tool means editing one line and the two cannot drift.

Version lookups never fail the run: each falls back to "unknown".

Verified with nextflow -profile stub -stub-run: 15 processes, 23 tasks,
exit 0. Cache safety confirmed by establishing a baseline without these
processes and resuming with them — 12 of the 13 existing processes
reported cached; only the two new ones and MULTIQC ran.

Also adds test/data/ to .gitignore (regenerated by make_stub_inputs.sh)
and a CHANGES.md entry for the hardening work since v1.0.0.
MISER_VERSION is the only process that instantiates the MisER image
purely to read version strings. On a resumed run every MisER task is
cached, so the image is never materialised — meaning this process was
the first in months to trigger a pull, and an environmental failure
there took the whole pipeline down.

errorStrategy 'ignore' on both version processes. Verified by forcing
MISER_VERSION to exit 1: Nextflow exits 0, MULTIQC still produces its
report, and software_versions.yml is simply absent.

Note the coupling: SOFTWARE_VERSIONS takes MISER_VERSION's output as
input, so if the latter is skipped there is no versions section at all.
The Installation block could not be copy-pasted: the clone URL and the
singularity pull commands were both template placeholders, and the stated
Nextflow requirement was still >=23.10.0 after the resourceLimits
migration raised the manifest floor to >=24.04.0. The version claim was
the damaging one — a user on 23.10 would believe they were supported.

Also notes SINGULARITY_TMPDIR, since a pull run from a head node fails
when it points at node-local scratch that only exists on compute nodes.
Checks the repository against nf-core conventions in CI rather than
expecting anyone to install nf-core tools locally. The exclusions in
.nf-core.yml are what keep this meaningful rather than noisy: LoReRNA
follows nf-core structure but is not hosted under the nf-core
organisation, so their CI, branding and release infrastructure checks
do not apply.
nf-core lint crashed before running any check with 'not enough values to
unpack'. Cause is in nf_core/utils.py, which does

    self.pipeline_prefix, self.pipeline_name = manifest.name.split('/')

so a manifest.name without a slash fails the unpack. 'lorerna' becomes
'dariusng28-collab/LoReRNA', which is also the form Nextflow resolves for
nextflow run <org>/<repo>. Nothing in the pipeline reads manifest.name and
it is not part of the task hash.

Also pins nf-core==4.1.0 in CI. It was installed unpinned, so a new
nf-core/tools release could fail CI on a commit that changed nothing —
the same reason NXF_VER is pinned.
The nine lint failures were all identity checks — nf-core branding
artwork, an nf-core org CI workflow, a name beginning nf-core/, a
homePage under github.com/nf-core, and registration with nf-core/configs.
Passing them would require renaming the pipeline and rehosting it in the
nf-core organisation, which is not what this repository is.

142 checks pass; none of the nine concerns pipeline behaviour.

Also corrects the logo exclusions, which were lowercase and therefore
never matched: the expected filenames derive from manifest.name, whose
short_name is LoReRNA.
Nextflow pulls both images automatically, so this needs no action on a
normal system. It fails when SINGULARITY_TMPDIR points at a path absent
on the machine running the pull — typically a cluster setting it to
node-local scratch, which exists on compute nodes but not the head node.

Covers the three failure modes seen in practice: the missing build
directory, pull timeouts when several tasks race for an uncached image,
and disk exhaustion during layer unpacking. Previously this was only
noted under Installation, not where anyone looks when it breaks.
Singularity unpacks image layers into scratch resolved as
SINGULARITY_TMPDIR -> TMPDIR -> /tmp. Schedulers routinely point TMPDIR
at node-local scratch that exists on compute nodes but not on the login
node where nextflow run performs the pull, so the first container pull
fails with an internal error about a build parent dir that says nothing
about the cause.

main.nf now checks this at startup, before any task is submitted, and
reports it in terms the user can act on. Verified three ways: silent
with no container engine, fires with a broken path, silent with a valid
one. The variable cannot be set from inside the pipeline, since the JVM
cannot alter the environment its child processes inherit, so detecting
and explaining is the strongest option available.

Also repairs validate_preflight.sh, which has been dead since the
nf-core layout refactor: it grepped main.nf for statements that moved to
workflows/lorerna.nf, and under set -e a grep with no match aborted the
script at check 1. Checks now search both files, two stale checks are
repointed at conf/modules.config and miser_qc_merge.nf, and a new check
covers the container runtime environment. 46 pass, 0 fail.
SGE sends SIGUSR1 as a resource-limit warning, so a job approaching its
memory or runtime ceiling exits 138. That code was absent from the retry
list while 137, 139 and 140 were all present, so instead of retrying with
scaled resources the task failed and cancelled the run — which is what
happened to ISOQUANT (Control3) on a cold validation run.

Applied to base, sge, slurm and lsf. A site config passed with -c that
defines its own errorStrategy overrides this and must be patched
separately.
Adds a Structure tab that draws the exon layout of every transcript at the
selected gene's locus, from gffcmp.combined.gtf. That file shares the TCONS_
namespace with the swish tables, so the join needs no mapping layer.

Introns are compressed, never stretched: at a typical locus only about a tenth
of the span is exonic, so on raw coordinates most exons render sub-pixel and
structural differences are invisible.

Optionally reads a bgzip-compressed, tabix-indexed reference GTF, queried by
region so the annotation is never parsed in full. With one loaded the
comparison baseline becomes the annotation's MANE Select transcript rather than
one of the assembled isoforms, and undetected reference models are drawn for
context. Rsamtools is a soft dependency; without it the reference controls hide
and the rest is unchanged.

Also documents that the DTU log2FC column is not a fold change. swish applies a
pseudocount of 5 to isoform proportions, confining the statistic to +/-0.263;
the panel reports the recovered usage shift in percentage points instead.

Verified against real pipeline output: exon coordinates checked through two
separately written parsers across 500 genes and 21,795 exons, and structural
invariants swept across all 35,762 gene keys.
@dariusng28-collab
dariusng28-collab merged commit b756b66 into main Aug 6, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant