Attention with a bias

17/01/2026 13 min

Listen "Attention with a bias"

Descargar episodio Ver en sitio original

Episode Synopsis

We review why some transformer models use a bias in attention and how ALiBi helps with long context. The provided sources focus on significant advancements in computational biology, specifically the evolution of the AlphaFold series for predicting 3D biomolecular structures. AlphaFold 2 revolutionized the field by using the Evoformer and attention mechanisms to interpret evolutionary and geometric data with near-experimental accuracy. Building on this, AlphaFold 3 expanded capabilities to include complexes with ligands and nucleic acids using an atom-level diffusion module. To further refine these models, HelixFold-S1 introduces a contact-guided sampling strategy that prioritizes likely binding sites to improve structural diversity and accuracy. Additionally, technical papers describe architectural components like ALiBi for handling long sequences and Swin Transformer's shifted windows. Together, these texts illustrate a shift toward more efficient, targeted sampling and integrated deep learning frameworks for complex molecular modeling.Sources:August 2021Swin Transformer: Hierarchical Vision Transformer using Shifted Windowshttps://arxiv.org/pdf/2103.14030April 2022 - (ALiBi)TRAIN SHORT, TEST LONG: ATTENTION WITH LINEAR
BIASES ENABLES INPUT LENGTH EXTRAPOLATIONhttps://arxiv.org/pdf/2108.12409April 2022:Swin Transformer V2: Scaling Up Capacity and Resolutionhttps://arxiv.org/pdf/2111.09883

More episodes of the podcast AI: post transformers

Squisher: Approximating the Fisher Information Matrix and use cases 17/01/2026

NVIDIA: TTT-E2E: Unlocking Long-Context Learning via End-to-End Test-Time Training 17/01/2026

Scaling laws: long context length and in context learning 17/01/2026

DeepSeek Engram: Scaling Large Language Models via Conditional Memory Lookup 14/01/2026

PageANN: Scalable Disk ANNS with Page-Aligned Graphs 07/12/2025

NeurIPS 2025: Homogeneous Keys, Heterogeneous Values 04/12/2025

NeurIPS 2025: Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free 29/11/2025

NeurIPS 2025: Large Language Diffusion Models 29/11/2025

NeurIPS 2025: Reinforcement Learning for Reasoning in Large Language Models with One Training Example 29/11/2025

NeurIPS 2025: Parallel Scaling Law for Language Models 29/11/2025

Ver todos los episodios

ZARZA We are Zarza, the prestigious firm behind major projects in information technology.

Attention with a bias

Listen "Attention with a bias"

Episode Synopsis

More episodes of the podcast AI: post transformers

WWW. Is it obsolete or not? Should we use it?

Do you work sitting down? Do active breaks

Bandwidth: Broadband or Narrowband?

Personnel recruitment via Web

Deep web or Invisible Internet

Subdomains, a glance with the experts!

Free Internet, a prediction in Nostradamus style

Educational Technology: From traditional to digital

Localhost, there’s no place like 127.0.0.1

Googling with breathtaking tricks you ignore

Internet Predators on the prowl

Gray Hat Hacking, those with ambiguous ethics…

Dot COM: The Internet’s dominant TLD