A hands-on machine learning and deep learning laboratory.
Understand the mathematics. Build the models. Reproduce the papers. Build the framework.
TorchLab is my personal, open-source laboratory for understanding machine learning and deep learning by implementing them.
The goal is to go beyond using existing frameworks or copying model architectures. I want to understand how learning algorithms work mathematically, how PyTorch implements them, how modern research architectures are constructed, and ultimately how to build the underlying deep learning framework myself.
The repository follows six connected stages:
- Machine learning: Study classical ML algorithms and implement them from scratch using NumPy.
- Deep learning theory: Derive the mathematics behind neural networks, optimization, backpropagation, and modern architectures.
- PyTorch: Understand tensors, automatic differentiation, model development, training, and the framework's internals.
- Research implementations: Read and reproduce influential deep learning papers using PyTorch.
- Torch from scratch: Build a minimal deep learning framework, from automatic differentiation to neural network layers and optimizers.
- Research on our framework: Reimplement the previously studied architectures using the custom framework and compare their behavior against PyTorch.
The repository is a work in progress. It currently contains several research implementations, a growing paper catalog, learning roadmaps, and the first component of the custom framework. The six stages describe both what is available today and where the laboratory is heading.
TORCHLAB
│
▼
┌───────────────────────┐
│ 01. Machine Learning │
│ Mathematics + NumPy │
└───────────┬───────────┘
▼
┌───────────────────────┐
│ 02. Deep Learning │
│ Theory & Derivations │
└───────────┬───────────┘
▼
┌───────────────────────┐
│ 03. PyTorch │
│ Framework & Training │
└───────────┬───────────┘
▼
┌───────────────────────┐
│ 04. Research Papers │
│ PyTorch Reproductions │
└───────────┬───────────┘
▼
┌───────────────────────┐
│ 05. Torch From Scratch│
│ Our Own Framework │
└───────────┬───────────┘
▼
┌───────────────────────┐
│ 06. Research Papers │
│ On Our Own Framework │
└───────────────────────┘
| Stage | Focus | Current status |
|---|---|---|
| 01 · Machine Learning | Classical ML, mathematical foundations, and NumPy implementations | Syllabus |
| 02 · Deep Learning Theory | Neural networks, backpropagation, optimization, and architecture derivations | Syllabus |
| 03 · PyTorch | PyTorch fundamentals, training, performance, and internals | Syllabus |
| 04 · Papers in PyTorch | Paper explanations, implementations, experiments, and a research catalog | Active |
| 05 · Torch From Scratch | Custom autograd, tensors, neural network layers, optimizers, and backends | In progress |
| 06 · Papers in Torch From Scratch | Reproducing research architectures with the custom framework | Planned |
Understanding the algorithms before using deep learning frameworks.
This stage is dedicated to classical machine learning, starting with mathematical intuition and progressing toward implementations using NumPy.
The machine learning syllabus covers:
- Linear and logistic regression
- Gradient descent and optimization
- Regularization, generalization, and model evaluation
- K-nearest neighbors, Naive Bayes, and support vector machines
- Decision trees, random forests, and ensemble methods
- K-means, Gaussian mixtures, and principal component analysis
- Probability and information theory
The intended structure for each topic is a mathematical explanation, an implementation, and experiments on small datasets.
Status: The syllabus is available. Individual topic implementations are planned.
Understanding what happens inside a neural network.
This stage focuses on mathematical derivations, computational graphs, and the theoretical foundations that connect classical neural networks to modern architectures.
Topics in the deep learning theory syllabus include:
- Perceptrons, multilayer perceptrons, and activation functions
- Backpropagation and the chain rule
- Optimization, initialization, and normalization
- Convolutional and recurrent neural networks
- Attention mechanisms and Transformers
- Tokenization and embeddings
- VAEs, GANs, and diffusion models
- Reinforcement learning and preference optimization
- Model scaling and efficiency
The goal is to understand not only how these methods work, but also why their design choices matter.
Status: The theoretical learning roadmap is available. Dedicated topic modules are planned.
Learning the framework that powers the research implementations.
This stage studies PyTorch from fundamental tensor operations to the systems used for training and inference.
The PyTorch syllabus covers tensors, autograd, nn.Module, data pipelines, training loops, mixed precision, profiling, distributed training, custom CUDA/Triton kernels, and model inference.
The objective is twofold: use PyTorch effectively for research, and understand its abstractions well enough to rebuild a smaller version in Stage 05.
Status: The syllabus is available. Dedicated framework tutorials and examples are planned.
From a published paper to an understandable, working implementation.
This is the main implementation-focused section of TorchLab.
It combines explanations of research ideas with PyTorch code, model architectures, and practical examples.
Beyond the existing packages, TorchLab maintains a larger reading and implementation catalog.
The catalog covers foundational neural networks, optimization, computer vision, Transformers, generative models, speech, and reinforcement learning.
It includes papers collected from research and educational resources, including Geoffrey Hinton's publications and the work covered by Umar Jamil, Priyam Mazumdar, and Aladdin Persson.
Important: Papers in the catalog are research and implementation candidates. Inclusion does not mean a paper has already been reproduced.
Browse the full research paper catalog →
Building the abstractions behind deep learning frameworks, one component at a time.
Using PyTorch is one thing. Understanding how automatic differentiation, tensors, neural network layers, and optimizers are implemented is another.
The goal of this stage is to develop a small, educational, PyTorch-inspired framework that can eventually run the architectures implemented in Stage 04.
The first available component is:
Forward-mode automatic differentiation with dual numbers
This introduces dual-number arithmetic and tangent propagation as a foundation for understanding automatic differentiation.
| Component | Scope | Status |
|---|---|---|
| Forward-mode autograd | Dual numbers and tangent propagation | Done |
| Reverse-mode autograd | Computational graphs, backward passes, and VJPs | Planned |
| Tensor engine | N-dimensional storage, broadcasting, views, and matrix operations | Planned |
| Tensor autograd | Differentiation over tensor operations | Planned |
| Neural network API | Modules, parameters, layers, and activations | Planned |
| Loss functions | MSE, cross-entropy, and numerical stability | Planned |
| Optimizers | SGD, momentum, Adam, and AdamW | Planned |
| Data utilities | Datasets, batching, and data loading | Planned |
| Attention | Scaled dot-product attention, multi-head attention, and caching | Planned |
| Backends | NumPy backend, followed by GPU exploration | Planned |
| Testing | Numerical gradient checks and PyTorch comparisons | Planned |
As the components mature, the intention is to assemble them into a reusable, importable framework.
This is an educational framework under development, not a replacement for production PyTorch.
Closing the loop: implementing research architectures using the framework we built ourselves.
The final stage aims to reproduce selected Stage 04 implementations without relying on PyTorch for their core model operations.
The intended progression is:
- Build the required tensor, differentiation, and neural network primitives.
- Reimplement the selected model using the custom framework.
- Validate numerical behavior and gradients against PyTorch.
- Compare training behavior, performance, and implementation complexity.
The initial direction is to revisit architectures already studied in the laboratory, beginning with their fundamental components.
Status: Roadmap. This stage depends on the development of the Stage 05 framework.
TorchLab/
│
├── 01-machine-learning/
│ └── README.md Classical ML syllabus
│
├── 02-deep-learning-theory/
│ └── README.md Deep learning theory syllabus
│
├── 03-pytorch/
│ └── README.md PyTorch learning syllabus
│
├── 04-papers-in-pytorch/
│ ├── README.md Research catalog and package status
│ ├── transformer/
│ │ ├── transformer/ Transformer model and notebooks
│ │ └── translator/ English → Arabic translation
│ ├── llama/
│ │ ├── x_llama/ Llama-style model and inference
│ │ └── models/ Attention and rotary embeddings
│ ├── diffusion/ Diffusion and Stable Diffusion
│ └── rlhf/
│ └── motion/ Motion Canvas animation source
│
├── 05-torch-from-scratch/
│ ├── README.md Custom framework roadmap
│ └── autograd/
│ └── forward-mode-dual-numbers/
│
├── 06-papers-in-torch-from-scratch/
│ └── README.md Future research implementations
│
├── website/ Vue + Vite website
├── assets/ TorchLab branding
└── README.md
The tree highlights the main learning and implementation areas; individual packages contain additional source files, notebooks, and assets.
Clone the repository:
git clone https://github.com/Esmail-ibraheem/TorchLab.git
cd TorchLabStart with the Transformer implementation:
cd 04-papers-in-pytorch/transformerRead the package's README.md to understand the architecture, then explore the model code, notebooks, and translation example.
Other available research packages can be found under 04-papers-in-pytorch/.
Return to the repository root and open:
cd 05-torch-from-scratch/autograd/forward-mode-dual-numbersThis contains the first automatic differentiation component.
Installation and execution details may differ between packages. Consult each package's documentation and dependency files where available.
TorchLab is an evolving learning and research project.
The next areas of development are:
- Expand classical machine learning implementations in NumPy.
- Add detailed deep learning derivations and numerical experiments.
- Develop the PyTorch learning modules.
- Expand the research paper implementation catalog.
- Add PPO and DPO fine-tuning implementations to the RLHF package.
- Implement forward-mode automatic differentiation with dual numbers.
- Implement reverse-mode automatic differentiation.
- Build the tensor engine and neural network API.
- Add optimizers, data utilities, and tests.
- Assemble the components into an importable deep learning framework.
- Reimplement selected research architectures on that framework.
The objective is not simply to collect more models, but to build a connected body of theory, implementations, experiments, and framework internals.
TorchLab is a personal learning laboratory, but contributions, suggestions, corrections, and discussions are welcome.
Potential contributions include improving mathematical explanations, identifying bugs in existing implementations, adding tests, proposing research papers, or extending the educational material.
If you find an issue or have an idea, feel free to open an issue or pull request.
TorchLab draws on research papers, open-source implementations, and educational resources.
Some foundational resources include:
- PyTorch documentation
- Deep Learning — Goodfellow, Bengio & Courville
- Dive into Deep Learning
- Andrej Karpathy's micrograd
- MiniTorch
- Attention Is All You Need
- Llama 2
- Denoising Diffusion Probabilistic Models
Individual packages include their own paper references and implementation details.
TorchLab — Understand it. Implement it. Rebuild it.
Built by Esmail Gumaan