Skip to content
Advanced45 lessons

AI Research Scientist

Learn to turn open questions into reproducible studies across model training, reinforcement learning, agents, and evaluation.

Aspiring research scientists and research engineers who need rigorous experimental judgment, implementation depth, and public evidence.

You can frame a falsifiable question, reproduce a baseline, run controlled ablations, evaluate agent and model behavior, and publish a defensible research artifact.

  1. 1Linear Algebra for MLFind hidden directions in a support-incident matrix with SVD, then use rank, PCA, truncation, and condition numbers without losing sight of what the numbers mean.Math & StatisticsEasy21 min
  2. 2Adam, Momentum, SchedulersTrace SGD, momentum, RMSProp, Adam, AdamW, schedules, and gradient clipping on a 100-to-1 loss valley. Learn what each optimizer buffer measures and how to validate a training choice.Math & StatisticsEasy23 min
  3. 3Probability for Machine LearningUse one API abuse-risk detector to learn events, random variables, distributions, conditional probability, independence, Bayes rule, and base-rate mistakes.Math & StatisticsEasy23 min
  4. 4Statistics and UncertaintyEstimate a review queue's abuse rate with bootstrap and Wilson intervals, distinguish sampling variation from bias, and assess calibration without overstating the evidence.Math & StatisticsEasy21 min
  5. 5Distributions and SamplingMatch binary outcomes, routes, tool-call counts, and latency to first distributions, then reject a simulation that doesn't fit the traces.Math & StatisticsEasy34 min
  6. 6Hypothesis Tests, Intervals, and pass@kCompare a code-generation model with paired evidence, uncertainty for lift, and pass@k under a fixed sampling budget.Math & StatisticsEasy26 min
  7. 7Validation and LeakageSplit access-review requests by time and user, block post-decision fields, fit preprocessing on training rows only, and treat public LLM benchmarks as contamination-prone.ML Algorithms & EvaluationMedium32 min
  8. 8Experiment Design and A/B TestingDesign a trustworthy online experiment for an incident-assistant change: randomize incidents, measure useful outcomes, quantify uncertainty, and reject false wins.ML Algorithms & EvaluationMedium41 min
  9. 9Reinforcement Learning BasicsTurn the earlier one-shot access-review label into an MDP. Compute discounted returns and Bellman backups, run value iteration and Q-learning, watch abandonment reverse a policy, and connect REINFORCE to LLM post-training.ML Algorithms & EvaluationMedium31 min
  10. 10Dataset Pipelines and Data QualityBuild versioned AI datasets with schema gates, grouped splits, contamination checks, and auditable receipts.ML Algorithms & EvaluationMedium30 min
  11. 11Experiment Tracking with MLflow and W&BRecord the inputs, results, artifacts, and limits behind a candidate decision, then verify them with local MLflow and W&B examples.Applied LLM EngineeringMedium31 min
  12. 12PyTorch Training LoopsTrace logits, gradients, updates, and validation in runnable PyTorch loops. Work through unequal micro-batches, CUDA AMP ordering, independent snapshots, and checkpoint resumption.ML Algorithms & EvaluationMedium36 min
  13. 13CUDA for ML TrainingFollow one access-ticket batch from CPU memory into CUDA kernels. Learn thread and memory hierarchy, coalesced access, roofline decisions, safe device placement, honest timing, and first-line diagnosis for setup, OOM, and throughput failures.Computing FoundationsEasy25 min
  14. 14The Transformer Architecture End-to-EndTrace a three-token prompt through causal attention, residual updates, and a vocabulary head. Build and test one complete decoder in PyTorch.Preparation & PrerequisitesEasy26 min
  15. 15Language Modeling & Next TokensLearn how next-token prediction becomes a trainable language model, from bigram counts and neural n-grams to causal Transformer generation and KV-cache serving.Preparation & PrerequisitesEasy29 min
  16. 16Scaled Dot-Product AttentionBuild scaled dot-product attention from a token sequence: Q/K/V routing, variance scaling, masks, multi-head shapes, KV-cache cost, and FlashAttention.Transformer Deep DivesHard64 min
  17. 17Layer Normalization: Pre-LN vs Post-LNUnderstand LayerNorm mechanics, Pre-LN versus Post-LN placement, RMSNorm simplification, gradient stability, and hybrid normalization layouts for deep transformers.Transformer Deep DivesHard41 min
  18. 18Inference: TTFT, TPS & KV CacheMap prefill vs decode bottlenecks, measure TTFT and decode cadence, and size KV cache so concurrent sequences fit on one GPU.Inference & Production ScaleHard44 min
  19. 19FlashAttention & Memory EfficiencyUnderstand how FlashAttention cuts auxiliary attention memory from O(n²) to O(n) with tiling and online softmax, and analyze its IO complexity.Inference & Production ScaleHard51 min
  20. 20Scaling Laws & Compute-Optimal TrainingLearn how Kaplan, Chinchilla, and inference-aware fits split a training budget across parameters and tokens, and when a smaller over-trained model wins on lifetime cost.Advanced Training & AdaptationHard42 min
  21. 21Pre-training Data at ScaleUnderstand how web-scale pre-training data is extracted, filtered, deduplicated, mixed, tokenized, and packed into training-ready shards, including decontamination, late-stage annealing, and synthetic-data tradeoffs.Advanced Training & AdaptationHard42 min
  22. 22Build GPT from Scratch LabBuild and train a tiny GPT end to end on Shakespeare: tokenize with GPT-style subwords, remap active token IDs, run causal self-attention, track validation loss, save a checkpoint, and sample text.Advanced Training & AdaptationHard35 min
  23. 23JAX for PyTorch ResearchersRead and modify JAX research code after the PyTorch GPT lab by making state, randomness, transformations, compilation, and timing explicit.Advanced Training & AdaptationHard45 min
  24. 24Instruction Tuning & Chat TemplatesTeach a base language model to answer as an assistant: curate grounded SFT rows, serialize chat turns exactly, choose loss targets, pack safely, and detect serving-time template drift.Core LLM FoundationsMedium32 min
  25. 25Continued Pretraining for Domain ShiftLearn when to keep the causal language-modeling objective and continue pretraining on domain text instead of jumping straight to SFT, and how to evaluate the trade-off against forgetting, cost, and downstream gain.Advanced Training & AdaptationHard38 min
  26. 26LLM-as-a-Judge EvaluationBuild rubric-based judges, detect unstable preferences, and measure agreement without mistaking a plausible score for ground truth.Applied LLM EngineeringMedium25 min
  27. 27Synthetic Data PipelinesBuild post-training synthetic data as a gated pipeline: Self-Instruct, Evol-Instruct, grounded execution, calibrated judges, preference pairs, diversity, decontamination, and versioned shards.Advanced Training & AdaptationHard34 min
  28. 28Supervised Fine-Tuning PipelineRun supervised fine-tuning as a real training system: choose the learning objective before the update surface, verify response-token loss and packing, track the real batch budget, save resumable checkpoints, and export on held-out behavior.Advanced Training & AdaptationHard38 min
  29. 29Mixed Precision TrainingExplore disappearing gradients, asymmetric rounding, and loss scaling. Test why ordinary AMP keeps FP32 updates before comparing training speed and memory.Advanced Training & AdaptationHard44 min
  30. 30Distributed Training: FSDP & ZeROMaster ZeRO stages, FSDP2 fully_shard architecture, 4D parallelism topology (DP, TP, PP, CP), and NCCL collective communication trade-offs.Advanced Training & AdaptationHard65 min
  31. 31Training Run OperationsRecover training state after interruption, budget checkpoint time, preserve effective batch and scheduler progress, and choose an adaptation recipe that fits the data and memory.Advanced Training & AdaptationHard31 min
  32. 32Data Labeling and Human FeedbackTurn traces into reviewed feedback data: preserve privacy, measure agreement, select useful cases, and protect independent evaluation.Applied LLM EngineeringMedium31 min
  33. 33Reward Modeling from Preference DataTrain reward models as a first-class post-training stage: validate chosen/rejected pairs and splits, fit a scalar reward head with Bradley-Terry loss, audit generalization, and decide when explicit rewards are worth the extra complexity.Advanced Training & AdaptationHard41 min
  34. 34RLHF & DPO AlignmentFollow preference labels into PPO-style RLHF or DPO. Calculate KL-shaped rewards, clipped policy updates, and response-only DPO gradients, then check whether better training scores improve held-out behavior.Advanced Training & AdaptationHard56 min
  35. 35Constitutional AI & Red TeamingTrace Constitutional AI from critique and revision to AI preference labels, then build honest red-team evaluations that separate model failures, false refusals, and judge errors.Advanced Training & AdaptationHard33 min
  36. 36RLVR & Verifiable RewardsUnderstand RLVR, a post-training approach that uses programmatic verification instead of learned human-preference rewards to improve checked outcomes in math, code, and other contract-driven tasks.Advanced Training & AdaptationHard43 min
  37. 37Function Calling & Tool UseBuild a safe tool-calling runtime that validates model requests, executes controlled actions, feeds observations back, and evaluates complete workflows.Applied LLM EngineeringMedium29 min
  38. 38ReAct & Plan-and-ExecuteCompare ReAct for tightly coupled tool use with Plan-and-Execute for longer workflows with explicit planning and replanning.Advanced Agents & RetrievalHard66 min
  39. 39Evaluating AI AgentsEvaluate model-promotion agent runs by final state, observable trace, safety gates, cost, and repeatability, then map private tests to public benchmarks.Applied LLM EngineeringMedium32 min
  40. 40LLM Benchmarks & LimitationsBuild an evaluation suite for a policy-answering LLM: score evidence use, understand public benchmark contracts, control judge bias, and make release decisions from private tests.Core LLM FoundationsMedium34 min
  41. 41KV Cache & PagedAttentionCalculate KV cache capacity, trace paged block allocation, and separate memory packing from prefix reuse and scheduling tradeoffs.Inference & Production ScaleHard48 min
  42. 42Continuous Batching & SchedulingTrace iteration-level slot reuse, budget prefill and decode work, and evaluate chunking or phase separation against TTFT, TPOT, and inter-token latency targets.Inference & Production ScaleHard44 min
  43. 43A/B Testing for LLMsTake one docs-assistant prompt duel from a golden-set rubric to a live resolution-rate test with sticky routing and registered guardrails.Inference & Production ScaleHard61 min
  44. 44Mechanistic InterpretabilityLearn how sparse autoencoders decompose transformer activations into candidate interpretable features, support circuit tracing, and enable controlled activation-steering experiments.Transformer Deep DivesHard46 min
  45. 45Capstone: Reproducible ML StudyTest a reward-shaping claim with paired trials, inspect regressions, and package the code, uncertainty, and evidence for review.Portfolio CapstonesHard52 min