Apple researchers have unveiled SimpleDesign, an innovative protein design model that bypasses traditional multi-stage training by learning directly from raw data. This new approach allows the model to simultaneously generate both the amino acid sequence and the three-dimensional structure of proteins, streamlining the design process.
A Simplified Approach to Protein Design
Published as a preprint on arXiv, SimpleDesign builds upon Apple’s earlier SimpleFold project. While SimpleFold focused on predicting protein structures from sequences using flow matching models, SimpleDesign extends this simplified architecture to protein design. It aims to create not only the final structure but also the specific amino acid sequence that will fold into that structure.
Unlike many existing models that require intermediate steps like converting protein structures into discrete latent representations before training generative models, SimpleDesign adopts an end-to-end training approach. It directly utilizes amino acid sequences and continuous 3D coordinates, eliminating the need for intermediate data transformations.
Training with a Massive Dataset
The research team trained SimpleDesign on a substantial dataset comprising over two million protein sequence-structure pairs. This data was primarily sourced from the AFESM dataset, which integrates predicted structures from the AlphaFold database and other sources. During training, the model engages in a process of controlled data corruption.
Specifically, parts of the amino acid sequences are randomly masked, and noise is introduced into the corresponding 3D structures. By varying the degree of disruption to both sequence and structure, the model learns to tackle different protein-related tasks:
- Protein Folding Analogy: When sequences are mostly intact but structures are significantly perturbed, the task resembles protein folding – predicting the structure from a known sequence.
- Inverse Folding Analogy: When structures are largely complete and sequences are heavily masked, it mimics inverse folding – generating an amino acid sequence that would form a specified structure.
- Co-design: When both sequences and structures are partially corrupted, the model learns to jointly process this information, enabling the synergistic design of protein sequences and their corresponding structures.
The study clarifies that protein folding is the process by which a chain of amino acids adopts a specific three-dimensional shape, while inverse folding involves finding a sequence that corresponds to a target structure.
Promising Results, Future Validation Needed
Initial results reported in the research paper indicate that SimpleDesign performs competitively across benchmark tests for protein co-design, structure generation, and sequence generation. The researchers also noted that the model can produce protein structures with plausible shapes, and the quality of the generated amino acid sequences is on par with or better than those produced by many competing multimodal models.
However, the current findings are based solely on computational evaluations. The research team has not yet experimentally validated whether the proteins designed by SimpleDesign can actually fold correctly, perform their intended functions, or operate safely within biological systems. Therefore, these results do not yet confirm the practical applicability of the proteins designed by the model.
The development of SimpleDesign represents a significant step forward in computational protein design, offering a more integrated and efficient approach. Further experimental validation will be crucial to determine its real-world impact in fields ranging from drug discovery to materials science.









