Peptone and NVIDIA introduce PepTron-o for ensemble-first structure prediction of the disordered proteome
Peptone and NVIDIA introduced PepTron-o, a sequence-to-ensemble model designed to represent the conformational diversity of intrinsically disordered proteins
Intrinsically disordered regions make up more than 30% of the human proteome, but leading structure predictors are primarily designed for well-folded domains. Those predictors commonly reduce an intrinsically disordered region to a single low-confidence representation, which cannot express the range of conformations the protein can adopt.
Peptone and NVIDIA developed PepTron-o to address that gap. PepTron-o is a sequence-to-ensemble model that learns the conformational diversity of disordered proteins directly and provides an automated way to improve existing ensembles.
Lowering the barrier to structural ensemble generation
- Peptone developed a synthetic dataset containing tens of thousands of proteins with intrinsically disordered regions. The proteins were generated using the Oppenheimer Product Engine and reweighted against high-quality experimental observables, including nuclear magnetic resonance chemical shifts.
- PepTron-o is built on NVIDIA BioNeMo. It uses multiple forms of parallelism, fused attention kernels, and mixed-precision arithmetic from large-scale pretraining through fine-tuning and inference.
- The approach reduces GPU requirements and makes this class of protein-structure generation model easier to deploy.
A general reweighting framework
Current generative models can produce structural ensembles, but those ensembles may mix physically plausible and implausible conformations. PepTron-o introduces a model-agnostic reweighting procedure that:
- evaluates each structure against high-quality physical observables;
- reweights the ensemble using experimental readouts and back-propagates the updated weights, allowing the base model to learn the correction signal; and
- improves subsequent ensemble generation.
Because the reweighting head is model-agnostic, it can be applied to other ensemble generators.
Measuring ensemble quality
Metrics designed for a single structure, such as RMSD and LDDT, do not adequately describe intrinsically disordered regions. Peptone and NVIDIA therefore proposed an ensemble-consistency metric that penalizes ensembles when they cannot be reweighted to satisfy experimental observations, or when satisfying those observations requires an extreme redistribution of weights.
Peptone CTO Carlo Fisicaro and NVIDIA HCLS Startups Developer Relations Lead EMEA Cedric Steenbeke were scheduled to present the work at VivaTech 2025 in Paris on Friday, June 13, at 11:20 CET. Their talk was titled “Redefining how we generate and evaluate multi-domain protein structural ensembles.”