A drop-in, hardware-agnostic library for Fused Triangle Multiplicative Updates across AlphaFold3 family models, powered by CuTe DSL. Uses roughly half the GPU memory with almost zero activation data at any sequence length. It is far faster for small sequences.
Framer AI AgentsDesign and publish professional sites with AI
Promoted
Maker
📌
The creators of AlphaFold won a Nobel Prize in 2024. I made its family of AI open source models faster with a CUDA kernel.
I created fast_trimul, a drop-in Python library for Fused Triangle Multiplicative Updates across AlphaFold3 family models.
It has a modular architecture, since it was designed to be easy to use in production, and it works on OpenFold3 and Protenix.
The CUDA kernel is written in the Python CuTe DSL, so the kernel codebase stays small and easy to adapt for other GPUs like H100 or B200.
GitHub: https://github.com/tiagomonteiro...
LinkedIn: https://lnkd.in/p/gFGnV3JA