Changelog
All notable changes to ProteoPy will be documented in this file. The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[Unreleased]
Added
Plotting (pr.pl)
peptide_intensities(),proteoform_intensities(): newfacet_byparameter splitting the samples (.obs) across a grid of subplots
Preprocessing (pr.pp)
summarize_peptides_by_neighbourhood_union(): collapses peptides that overlap in the protein sequence, keeping the most abundant member of each group. Reimplements CCprofiler’ssummarizeAlternativePeptideSequences(topN = 1)Tools (
pr.tl):peptide_proximity()tests whether the peptides of a proteoform cluster sit closer together in the protein sequence than a random grouping of the same size. A reimplementation of CCprofiler’sevaluateProteoformLocation, which is the characterisation COPF applies to the proteoform groups it detects.
Fixed
Datasets (pr.datasets) and Download (pr.download)
williams_2018()no longer conflates measured zeros with missing values. Three defects, each masked by another:zeros in
.Xwere coerced tonp.nan, discarding 13,547 genuine measurementscharge-state summation used pandas’ default
min_count=0, so a group with no measurements summed to0.0, inventing 3,324the same summation skipped
NaNinside a partially measured group, reporting a partial total as complete for 260 cells
This changes
.X, and therefore the output ofpr.download.williams_2018(). Verified against the PRIDE deposit used by the original publication: the two are now bit-identical, with an identical missingness patternwilliams_2018(): newzero_to_naparameter (defaultFalse), for consistency with sibling functions; mutually exclusive withfill_na. Governs zero semantics only, and does not restore the summation defects above
Changed
Preprocessing (pr.pp)
normalize_median():now defaults to
log_space=Truebatch_idparameter renamed togroup_byzeros_to_naparameter renamed tozero_to_na, for consistency with sibling functionsno longer accepts sparse
.Xinput
[0.1.1] - 2025-03-24
Added
Preprocessing (pr.pp)
summarize_modifications()for modification summarization
Analysis (pr.tl)
ANOVA support in
differential_abundance()
Plotting (pr.pl)
binary_heatmap(),box(),volcano(),peptides_on_sequence(),peptides_on_prot_sequence()print_statsparameter across multiple plot functions
Datasets (pr.datasets)
williams_2018()andkarayel_2020()download functions
Utilities (pr.utils)
public API:
is_proteodata(),check_proteodata(),is_log_transformed()
Documentation (docs/)
Sphinx documentation site
proteoform inference and protein-level analysis tutorials
Changed
Reader (pr.read)
diann()supports version >=1.9.1, with automatic version dispatch
Preprocessing (pr.pp)
impute_downshift(): now supportsgroup_bynormalize_median(): newmethodparameterremove_contaminants(): defaults toinplace=True
Validation (pr.utils)
is_proteodata()now checks for NaN in ID columns, infinite values in.X/layers, and obs/var index sync
Fixed
Plotting (pr.pl)
volcano_plot(): type incompatibility and label displayn_cat1_per_cat2_hist(): minimum bin width
[0.1.0] - 2025-01-29
Initial release of ProteoPy.
Added
Reader (pr.read)
DIA-NN and generic long-format tables
Annotation (pr.ann)
annotation of samples (
.obs) and variables (.var)
Preprocessing (pr.pp)
quality control: completeness filtering, CV calculation, contaminant removal
median normalization, downshift imputation
Analysis (pr.tl)
differential abundance: t-test, Welch’s test, ANOVA with multiple testing correction
proteoform inference: COPF algorithm reimplementation for detecting functional proteoform groups
Plotting (pr.pl)
volcano, abundance rank, intensity distribution, CV, correlation matrix and hierarchical clustering profile plots
Datasets (pr.datasets)
built-in example datasets (Karayel 2020)