Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion DESCRIPTION
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
Package: visPedigree
Type: Package
Title: Tidying, Analysis, and Fast Visualization of Animal and Plant Pedigrees
Version: 1.10.0
Version: 1.10.1
Authors@R: person("Sheng","Luan", email="luansheng@gmail.com", role = c("aut","cre"))
Description: Provides tools for the analysis and visualization of animal and
plant pedigrees. Analytical methods include equivalent complete generations,
Expand Down
8 changes: 8 additions & 0 deletions NEWS.md
Original file line number Diff line number Diff line change
@@ -1,3 +1,11 @@
# Changes in version 1.10.1

## Bug fixes
1. **`pedexport(software = "wombat")` now uses the integer layout**: The 1.10.0 character-ID WOMBAT format was rejected by the real WOMBAT program — the WOMBAT manual (section 6.3) requires three integer variables with codes in 0..2147483647, offspring codes numerically larger than either parent, and unknown parents coded `0`. The `wombat` format now shares the integer layout with `blupf90` (including the `xref` attribute and `<file>.xref` mapping file), satisfying these requirements by construction. Non-zero custom `missing` symbols are rejected for `wombat`. Verified end-to-end against the WOMBAT 26-05-2025 binary.

## New features
1. **`pedexport(software = "hiblup")`**: New format for HIBLUP's `--pedigree` file: character columns `animal`/`sire`/`dam`, no header by default, missing parents coded `"0"`. Verified end-to-end against the HIBLUP binary (estimated additive variance and EBVs identical to ASReml/WOMBAT on the same pedigree).

# Changes in version 1.10.0 released on 8 Aug 2026
## New features
1. **`pedexport()` for breeding software formats**: New function `pedexport()` converts a `tidyped` pedigree into the file format of common animal and plant breeding programs - BLUPF90, ASReml, Echidna, WOMBAT, MTDFREML, DMU - plus a generic `numeric` layout and an in-memory `sommer` format. Character formats (`asreml`, `echidna`, `wombat`, `sommer`) keep the original character IDs; numeric formats (`blupf90`, `mtdfreml`, `dmu`, `numeric`) renumber the pedigree and carry an `xref` mapping back to the original IDs (attribute on the returned table, and a `<file>.xref` file when `file` is given), mirroring RENUMF90's `_XrefID` file. Defaults follow each program's conventions for header, separator, and missing-parent symbol, and every format sorts rows so parents precede offspring.
Expand Down
92 changes: 61 additions & 31 deletions R/pedexport.R
Original file line number Diff line number Diff line change
Expand Up @@ -21,15 +21,23 @@
#' parents encoded as \code{"0"}. Rows are sorted by generation
#' (\code{Gen}) so founders appear first. Compatible with ASReml-R
#' and ASReml-SA (character IDs require the \code{!ALPHA} qualifier;
#' the header line requires the \code{!SKIP 1} qualifier).}
#' the header line requires the \code{!SKIP 1} qualifier, placed
#' immediately after the pedigree file name).}
#' \item{\code{"echidna"}}{Identical character layout and defaults to
#' \code{"asreml"}: columns \code{animal}, \code{sire}, and \code{dam},
#' a header by default, missing parents encoded as \code{"0"}, and rows
#' sorted by generation.}
#' \item{\code{"wombat"}}{Three character columns (\code{animal},
#' \code{sire}, \code{dam}), no header, missing parents encoded as
#' \code{"0"}. Wombat accepts alphanumeric IDs and recodes them
#' internally. Rows are sorted so parents precede offspring.}
#' \item{\code{"hiblup"}}{Three character columns (\code{animal},
#' \code{sire}, \code{dam}), no header by default, missing parents
#' encoded as \code{"0"}, and rows sorted by generation. Compatible
#' with HIBLUP's \code{--pedigree} file.}
#' \item{\code{"wombat"}}{Integer layout identical to \code{"blupf90"}
#' (\code{IndNum}, \code{SireNum}, \code{DamNum}), no header, missing
#' parents encoded as \code{0}. WOMBAT requires integer codes in the
#' range 0 to 2147483647, with every offspring numerically larger than
#' either parent and unknown parents coded \code{0} (WOMBAT manual
#' section 6.3); the \code{tidyped()} integer index satisfies both
#' ordering constraints by construction.}
#' \item{\code{"mtdfreml"}}{Identical integer layout to \code{"blupf90"}.
#' Compatible with MTDFREML and MTGSAM.}
#' \item{\code{"dmu"}}{Identical integer layout to \code{"blupf90"}.
Expand All @@ -51,27 +59,30 @@
#' @param sep Character scalar. Field separator used when writing to
#' \code{file}. Defaults to a single space (\code{" "}), which every
#' file-based format accepts. BLUPF90 requires a space; WOMBAT, MTDFREML,
#' and DMU accept spaces or TABs; ASReml and Echidna accept spaces, TABs,
#' or commas; and the generic \code{"numeric"} format accepts any
#' DMU, and HIBLUP accept spaces or TABs; ASReml and Echidna accept spaces,
#' TABs, or commas; and the generic \code{"numeric"} format accepts any
#' single-byte, non-empty separator supported by
#' \code{\link[data.table]{fwrite}}. ASReml/Echidna comma-delimited output
#' requires a \code{.csv} file extension or the \code{!CSV} qualifier.
#' @param header Logical scalar or \code{NULL}. Whether to include a column
#' header line. \code{NULL} (default) uses the software-specific default:
#' \code{TRUE} for \code{"asreml"}, \code{"echidna"}, and \code{"numeric"},
#' \code{FALSE} for \code{"blupf90"}, \code{"wombat"}, \code{"mtdfreml"}
#' and \code{"dmu"}. Ignored for \code{"sommer"}. Note for ASReml and
#' Echidna: a header line is read as data unless the \code{!SKIP 1}
#' qualifier is used in the command file.
#' \code{FALSE} for \code{"blupf90"}, \code{"wombat"}, \code{"mtdfreml"},
#' \code{"dmu"}, and \code{"hiblup"}. Ignored for \code{"sommer"}. Note
#' for ASReml and Echidna: a header line is read as data unless the
#' \code{!SKIP 1} qualifier is used in the command file (it must be placed
#' immediately after the pedigree file name — trailing \code{!SKIP} after
#' other qualifiers is silently ignored by ASReml-SA).
#' @param missing Character or integer scalar. Symbol for missing parents.
#' \code{NULL} (default) uses the software-specific default: \code{0L} for
#' numeric formats, \code{"0"} for \code{"asreml"}, \code{"echidna"}, and
#' \code{"wombat"}.
#' Numeric formats (\code{"blupf90"}, \code{"mtdfreml"}, \code{"dmu"},
#' \code{"numeric"}) require a single integer value; \code{"asreml"},
#' \code{"echidna"}, and \code{"wombat"} accept a character value (numeric
#' values are converted to character). Ignored for \code{"sommer"}, which
#' always codes missing parents as \code{NA}.
#' \code{"hiblup"}.
#' Numeric formats (\code{"blupf90"}, \code{"wombat"}, \code{"mtdfreml"},
#' \code{"dmu"}, \code{"numeric"}) require a single integer value, and for
#' \code{"wombat"} the value must be \code{0} (WOMBAT manual section 6.3);
#' \code{"asreml"}, \code{"echidna"}, and \code{"hiblup"} accept a
#' character value (numeric values are converted to character). Ignored
#' for \code{"sommer"}, which always codes missing parents as \code{NA}.
#'
#' @return A \code{data.table} in the target format, returned invisibly.
#' Numeric formats carry an \code{xref} attribute mapping each numeric ID
Expand All @@ -88,9 +99,9 @@
#' \tabular{llll}{
#' \strong{software} \tab \strong{Col 1} \tab \strong{Col 2} \tab
#' \strong{Col 3} \cr
#' blupf90 / mtdfreml / dmu / numeric \tab \code{IndNum} (integer) \tab
#' \code{SireNum} (integer) \tab \code{DamNum} (integer) \cr
#' asreml / echidna / wombat \tab \code{animal} (character) \tab
#' blupf90 / wombat / mtdfreml / dmu / numeric \tab \code{IndNum}
#' (integer) \tab \code{SireNum} (integer) \tab \code{DamNum} (integer) \cr
#' asreml / echidna / hiblup \tab \code{animal} (character) \tab
#' \code{sire} (character) \tab \code{dam} (character) \cr
#' sommer \tab \code{ID} (character) \tab \code{Sire} (character) \tab
#' \code{Dam} (character) \cr
Expand All @@ -107,8 +118,11 @@
#' \strong{missing} \tab \strong{IDs} \tab \strong{notes} \cr
#' blupf90 \tab no \tab spaces only \tab \code{0} \tab integer \tab
#' TAB separators rejected \cr
#' wombat \tab no \tab space / TAB \tab \code{"0"} \tab character \tab
#' alphanumeric accepted, recoded internally \cr
#' wombat \tab no \tab space / TAB \tab \code{0} \tab integer \tab
#' integer codes 0..2147483647; offspring code must exceed both parent
#' codes (manual section 6.3) \cr
#' hiblup \tab no \tab space / TAB \tab \code{"0"} \tab character \tab
#' alphanumeric accepted; missing parents may also be \code{NA} \cr
#' mtdfreml \tab no \tab space / TAB \tab \code{0} \tab integer \tab \cr
#' dmu \tab no \tab space / TAB \tab \code{0} \tab integer \tab \cr
#' asreml \tab yes \tab space / TAB / comma \tab \code{"0"} \tab
Expand All @@ -127,7 +141,7 @@
#' All numeric formats sort by \code{IndNum} ascending, which guarantees that
#' parent rows appear before offspring rows (parents always receive a smaller
#' integer index after \code{tidyped()} topological sorting). The
#' \code{"asreml"}, \code{"echidna"}, \code{"wombat"} and \code{"sommer"}
#' \code{"asreml"}, \code{"echidna"}, \code{"hiblup"} and \code{"sommer"}
#' formats sort by \code{Gen} ascending for the same reason.
#'
#' \strong{Optional \code{tidyped()} columns:}
Expand Down Expand Up @@ -182,6 +196,16 @@
#' out_echidna <- pedexport(tp, software = "echidna")
#' identical(out_echidna, out_asreml)
#'
#' # HIBLUP: character IDs, no header by default
#' out_hiblup <- pedexport(tp, software = "hiblup")
#' head(out_hiblup)
#'
#' # WOMBAT requires integer codes (manual section 6.3): same integer
#' # layout as BLUPF90, with the xref mapping back to original IDs
#' out_wombat <- pedexport(tp, software = "wombat")
#' head(out_wombat)
#' head(attr(out_wombat, "xref"))
#'
#' # Numeric formats carry the ID mapping back to character IDs
#' out_dmu <- pedexport(tp, software = "dmu")
#' head(attr(out_dmu, "xref"))
Expand All @@ -205,8 +229,9 @@
#' @import data.table
#' @export
pedexport <- function(ped,
software = c("blupf90", "asreml", "echidna", "wombat",
"mtdfreml", "dmu", "numeric", "sommer"),
software = c("blupf90", "asreml", "echidna", "hiblup",
"wombat", "mtdfreml", "dmu", "numeric",
"sommer"),
file = NULL,
sep = " ",
header = NULL,
Expand Down Expand Up @@ -239,7 +264,7 @@ pedexport <- function(ped,

# ---- 2. Software-specific defaults ----
asreml_like <- c("asreml", "echidna")
use_char <- software %in% c(asreml_like, "wombat", "sommer")
use_char <- software %in% c(asreml_like, "hiblup", "sommer")
def_header <- software %in% c(asreml_like, "numeric")
def_miss <- if (use_char) "0" else 0L

Expand All @@ -253,22 +278,27 @@ pedexport <- function(ped,
} else if (use_char) {
if ((!is.character(missing) && !is.numeric(missing)) ||
length(missing) != 1L || is.na(missing)) {
stop("For the 'asreml', 'echidna', and 'wombat' formats, 'missing' ",
stop("For the 'asreml', 'echidna', and 'hiblup' formats, 'missing' ",
"must be a single character (or numeric) value.", call. = FALSE)
}
missing <- as.character(missing)
if (!nzchar(missing)) {
stop("For the 'asreml', 'echidna', and 'wombat' formats, 'missing' ",
stop("For the 'asreml', 'echidna', and 'hiblup' formats, 'missing' ",
"must not be empty.", call. = FALSE)
}
} else {
missing_int <- suppressWarnings(as.integer(missing))
if (!is.numeric(missing) || length(missing) != 1L || is.na(missing) ||
is.na(missing_int) || missing_int != missing) {
stop("For numeric formats (blupf90, mtdfreml, dmu, numeric), ",
stop("For numeric formats (blupf90, wombat, mtdfreml, dmu, numeric), ",
"'missing' must be a single integer value.", call. = FALSE)
Comment on lines 290 to 294

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Require the WOMBAT missing-parent code to be zero

When software = "wombat" is combined with a custom integer such as missing = -99L, this generic numeric validation accepts it and .pedexport_num() replaces every unknown parent with that value. The resulting file violates WOMBAT's stated requirement that unknown parents be coded as 0, so an explicitly supported argument can produce an unusable or misinterpreted WOMBAT pedigree; reject nonzero missing values for this format.

Useful? React with 👍 / 👎.

}
missing <- missing_int
if (software == "wombat" && missing != 0L) {
stop("WOMBAT requires unknown parents to be coded 0 (manual section ",
"6.3); 'missing' must be 0 for software = \"wombat\".",
call. = FALSE)
}
}

# ---- 4. Build output table ----
Expand Down Expand Up @@ -406,7 +436,7 @@ pedexport <- function(ped,
stop("BLUPF90 pedigree files require sep = \" \".", call. = FALSE)
}

whitespace_formats <- c("wombat", "mtdfreml", "dmu")
whitespace_formats <- c("wombat", "mtdfreml", "dmu", "hiblup")
if (software %in% whitespace_formats && !(sep %in% c(" ", "\t"))) {
stop(sprintf("%s pedigree files require a space or TAB separator.",
toupper(software)), call. = FALSE)
Expand All @@ -431,7 +461,7 @@ pedexport <- function(ped,
#' @noRd
.validate_export_fields <- function(ped, software, missing, sep) {
fields <- ped$Ind
if (software %in% c("asreml", "echidna", "wombat")) {
if (software %in% c("asreml", "echidna", "hiblup")) {
fields <- c(fields, ped$Sire, ped$Dam, missing)
}
fields <- as.character(fields[!is.na(fields)])
Expand Down
Loading
Loading