Theoretical Foundations#
- 1.1 The Spin-Glass Analogy
- 1.2 The Boltzmann Distribution and Equilibrium
- Thermal Equilibrium and the Canonical Ensemble
- The Boltzmann Distribution
- Statistical Physics Foundations: Derivation from the Microcanonical Ensemble
- The Partition Function: A Computational Bottleneck
- The Thermodynamic Limit and Ensemble Equivalence
- The Boltzmann Distribution in Neural Networks
- The Meaning of Equilibrium in a Neural Context
- From Physics to Learning
- Summary
- 1.3 The Need for Noise: Escaping Spurious Minima
- The Hopfield Network: Deterministic Descent
- Memories as Energy Minima
- The Curse of Spurious Minima
- The Limits of Deterministic Search
- Enter Noise: Stochastic Dynamics and Thermal Fluctuations
- Statistical Physics Foundations: Noise as Thermal Equilibrium
- Simulated Annealing: From Exploration to Exploitation
- The Dual Role of Noise in Learning
- Summary: From Deterministic Traps to Stochastic Freedom
- 1.4 The Renormalization Group: From Microscopic Spins to Macroscopic Features
- 2.1 A Brief Recap: Linear Neurons and Limitations
- The Linear Neuron: Weighted Sum and Threshold
- Geometric Interpretation: Linear Separability
- The Perceptron Learning Algorithm
- The Limits of Linear Separability: The XOR Problem
- The Historical Impact: The First AI Winter
- From Feedforward Classification to Probabilistic Generation
- Connection to Energy-Based Models
- Summary
- 2.2 Recurrent Networks and Content-Addressable Memory
- From Feedforward to Recurrence
- The Hopfield Network: A Prototypical Recurrent Architecture
- Energy Landscape and Convergence
- Storing Memories as Attractors
- Retrieval Dynamics: Pattern Completion and Error Correction
- Content-Addressable vs. Address-Based Memory
- Capacity Limitations and Spurious Attractors
- From Hopfield to Boltzmann: The Missing Ingredient
- Summary
- 2.3 Hebbian Learning as Sculpting Energy
- 3.1 Defining the Objective: Low Energy for Real Data
- 3.2 The Intractable Partition Function Problem
- The Partition Function: Definition and Its Implications
- Why the Partition Function Matters
- The Partition Function as a Free Energy Barrier
- The Consequence: Exact Maximum Likelihood Is Impossible
- Hessian and Non-Convexity
- Approaches to Taming the Intractability
- The Partition Function in the Era of Deep Learning
- Summary
- 3.3 Contrastive Divergence
- The Core Insight: Truncated Markov Chains
- The CD-k Algorithm in Detail
- Geometric Interpretation: Approximating the Gradient
- Why Contrastive Divergence Works: A Heuristic Justification
- Limitations and Caveats
- From Contrastive Divergence to Modern Approximations
- Practical Implementation: Mini-Batches and Sampling
- Summary
- 4.1 The Classic Boltzmann Machine: Visible and Hidden Symmetry
- The Architecture: A Fully Connected Stochastic Network
- The Energy Function
- Stochastic Dynamics and Equilibrium
- Learning Objective: Maximum Likelihood with Hidden Variables
- The Wake-Sleep Algorithm: A Biological Metaphor
- The Computational Bottleneck: Equilibrium Sampling
- The Unconstrained Connectivity: A Double-Edged Sword
- The Path Forward: Restricted Architectures
- Recap: The Classic Boltzmann Machine
- Summary
- 4.2 Restricted Boltzmann Machine (RBM)
- The Architectural Restriction: A Bipartite Graph
- Energy Function and Joint Distribution
- Probability Distributions in RBMs
- The Key Computational Advantage: Conditional Independence
- Efficient Gibbs Sampling: Block Updates
- Learning the RBM: Contrastive Divergence Revisited
- Free Energy and Marginal Probability
- Free Energy and Its Analytical Form
- Free Energy vs. Expected Energy
- RBMs as Product of Experts
- Variants: Handling Different Data Types
- Limitations of the RBM
- The RBM’s Place in Deep Learning History
- Summary
- 4.3 Beyond Single Layers: Stacking for Deep Learning
- The Motivation: Hierarchical Feature Learning
- Deep Belief Networks: Composition of RBMs
- Greedy Layer-Wise Training
- Why Greedy Layer-Wise Training Works
- Deep Boltzmann Machines (DBMs)
- From Unsupervised Pre-Training to Discriminative Fine-Tuning
- The End of the Pre-Training Era and the Legacy of Stacked RBMs
- Stacking Beyond RBMs: The General Principle
- Summary
- Part I Summary: The Core Bottleneck and the Path Forward