Artificial intelligence (AI) has created new possibilities for designing tailor-made proteins to solve everything from medical to ecological problems. A research team has now successfully applied a computer-based natural language processing model to protein research. Completely independently, the ProtGPT2 model designs new proteins that are capable of stable folding and could take over defined functions in larger molecular contexts. The model and its potential are detailed scientifically in Nature Communications.
Natural languages and proteins are actually similar in structure. Amino acids arrange themselves in a multitude of combinations to form structures that have specific functions in the living organism – similar to the way words form sentences in different combinations that express certain facts.
In recent years, numerous approaches have therefore been developed to use principles and processes that control the computer-assisted processing of natural language in protein research.
"Natural language processing has made extraordinary progress thanks to new AI technologies. Today, models of language processing enable machines not only to understand meaningful sentences but also to generate them themselves. Such a model was the starting point of our research. With detailed information concerning about 50 million sequences of natural proteins, my colleague Noelia Ferruz trained the model and enabled it to generate protein sequences on its own. It now understands the language of proteins and can use it creatively. We have found that these creative designs follow the basic principles of natural proteins," says the senior author.
The language processing model transferred to protein evolution is called "ProtGPT2". It can now be used to design proteins that adopt stable structures through folding and are permanently functional in this state. In addition, the biochemists have found out, through complex investigations, that the model can even create proteins that do not occur in nature and have possibly never existed in the history of evolution. These findings shed light on the immeasurable world of possible proteins and open a door to designing them in novel and unexplored ways.
There is a further advantage: Most proteins that have been designed de novo so far have idealized structures. Before such structures can have a potential application, they usually must pass through an elaborate functionalization process – for example by inserting extensions and cavities – so that they can interact with their environment and take on precisely defined functions in larger system contexts. ProtGPT2, on the other hand, generates proteins that have such differentiated structures innately, and are thus already operational in their respective environments.
"Our new model is another impressive demonstration of the systemic affinity of protein design and natural language processing. Artificial intelligence opens up highly interesting and promising possibilities to use methods of language processing for the production of customized proteins. At the University of Bayreuth, we hope to contribute in this way to developing innovative solutions for biomedical, pharmaceutical, and ecological problems," says the senior author.
https://www.nature.com/articles/s41467-022-32007-7
http://sciencemission.com/site/index.php?page=news&type=view&id=publications%2Fprotgpt2-is-a-deep&filter=22
Artificial intelligence enables the design of novel proteins
- 1,168 views
- Added
Latest News
Protein that helps COVID-19…
By newseditor
Posted 26 Jul
Spinal Muscular Atrophy (SM…
By newseditor
Posted 26 Jul
Link between bowel movement…
By newseditor
Posted 26 Jul
Inhibition of IL-11 signall…
By newseditor
Posted 25 Jul
Brain changes linked to obe…
By newseditor
Posted 25 Jul
Other Top Stories
Burst of morning gene activity tells plants when to flower
Read more
How plants bind their green pigment chlorophyll
Read more
A topical gel to protect farmers against pesticide-induced neuronal…
Read more
Exploiting epigenetic variation for plant breeding
Read more
Plant-based toxin modified to target tumors
Read more
Protocols
A systems biology approach…
By newseditor
Posted 24 Jul
quantms: a cloud-based pipe…
By newseditor
Posted 22 Jul
Emerging tools and best pra…
By newseditor
Posted 19 Jul
Directly selecting cell-typ…
By newseditor
Posted 17 Jul
PUFFFIN: an ultra-bright, c…
By newseditor
Posted 16 Jul
Publications
Hepatocyte-intrinsic SMN de…
By newseditor
Posted 26 Jul
Aberrant bowel movement fre…
By newseditor
Posted 26 Jul
A pseudoautosomal glycosyla…
By newseditor
Posted 26 Jul
Microglia protect against a…
By newseditor
Posted 26 Jul
Rigor and reproducibility i…
By newseditor
Posted 26 Jul
Presentations
Myelin plasticity in the ve…
By newseditor
Posted 10 Jun
Hydrogels in Drug Delivery
By newseditor
Posted 12 Apr
Lipids
By newseditor
Posted 31 Dec
Cell biology of carbohydrat…
By newseditor
Posted 29 Nov
RNA interference (RNAi)
By newseditor
Posted 23 Oct
Posters
A chemical biology/modular…
By newseditor
Posted 22 Aug
Single-molecule covalent ma…
By newseditor
Posted 04 Jul
ASCO-2020-HEALTH SERVICES R…
By newseditor
Posted 23 Mar
ASCO-2020-HEAD AND NECK CANCER
By newseditor
Posted 23 Mar
ASCO-2020-GENITOURINARY CAN…
By newseditor
Posted 23 Mar