Interpretation and known limitations

Codon usage reflects interacting effects of mutation, selection, drift, gene expression, amino-acid composition, genome composition, recombination, population history, and data selection. A high host CAI or codon similarity does not directly demonstrate improved expression or viral fitness. Experimental validation and an appropriate biological null model are required.

Reference construction is often the largest source of avoidable bias. Use curated complete CDSs, document whether highly expressed genes were selected, avoid mixing genetic codes, preserve database releases, and compare sensitivity across reference definitions. Very short genes and compositionally unusual proteins provide few independent synonymous observations.

Optimization modifies nucleotide sequences and may affect RNA structure, splicing, transcription, innate immune recognition, synthesis, and regulatory motifs not included in the selected constraint set. CodonAdaptPy therefore returns multiple candidates and explicit constraint violations. Candidates require downstream in silico review and experimental validation.

Root-to-tip regression and deterministic lineage-through-time output are rapid temporal diagnostics. A fitted root date is sensitive to rooting, recombination, among-lineage rate variation, sampling bias, and weak temporal structure. These outputs are not substitutes for relaxed-clock posterior inference or Bayesian skyline effective-population-size estimates. Report R², the rooting procedure, sample-date coverage, excluded records, and the exact tree/alignment method.