How We Built LeetLLM
How LeetLLM keeps a 196-lesson visible curriculum coherent: Git-backed bundles, research packets, themed visual builds, bounded validators, and a no-traffic Cloud Run candidate before readers arrive.

Blog
Practical notes on LLM systems, evaluation, agents, inference, and developer workflows.
How LeetLLM keeps a 196-lesson visible curriculum coherent: Git-backed bundles, research packets, themed visual builds, bounded validators, and a no-traffic Cloud Run candidate before readers arrive.


Compare 2026 AI coding plan prices, usage meters, model access, and review surfaces across Cursor, Claude Code, Codex, Copilot, Devin, Gemini Code Assist, and secondary model plans.

Run Qwen3.6 locally with llama.cpp and Unsloth GGUF. Pin 27B or 35B-A3B, budget total memory from Unsloth's ladder, expose a local endpoint, then test MTP only after the plain run works.

Five AI engineering project ideas built around evidence a reviewer can inspect: eval rows, traces, tests, costs, failures, and design decisions.

Pick an AI engineering role, grow one incident-assistant project from tests to measured deployment, and collect evidence that matches that starting point.

Hosted DeepSeek V4, the open checkpoints, and the Hugging Face parameter badges are different artifacts. Pin the exact build before you route traffic.

Compare ChatGPT/Codex, MiniMax, Qwen Cloud, Z.AI, and direct OpenAI API billing for OpenClaw. Prices, quotas, setup, and fallback risks verified August 2026.

Run Gemma 4 locally with Ollama: pick an exact tag, leave memory for context, verify placement with ollama ps, and benchmark the native API.

Compare vLLM, SGLang, TensorRT-LLM, Ollama, and llama.cpp by deployment surface, checkpoint recipes, latency, throughput, and operating cost.

Fifty LLM engineering questions that connect model mechanics to serving, retrieval, agents, evaluation, and safety through evidence and failure diagnosis.

Compare Cursor, Codex, Copilot, and Claude Code through one authorization change. Choose by operating mode, review surface, cost meter, and team controls.

Compare AI engineering pay without mixing BLS wages, posted salary, and total compensation. Match role family, level, location, equity terms, and source date first.

Pick an LLM lane by data boundary and license first, then compare task quality, serving evidence, and cost only among the options you're allowed to run.

Million-token windows can hold a bounded corpus, but capacity doesn't guarantee reliable retrieval. Measure effective context, watch length-based prices, and choose between full context and RAG.

Build a working AI agent as a plain Python loop: the model proposes a tool, your runtime validates and executes it, and the observation drives the next turn.

Choose prompting, retrieval, fine-tuning, or a hybrid by tracing each failed eval case to missing evidence, unclear instructions, recurring behavior, or a hard application rule.

AI engineer isn't one job. Applied product work, ML platforms, and research engineering ship different artifacts. Live postings show how to tell them apart.

SWE-bench scores a system, not a model. Walk through the task format, harness, variants, Verified set, contamination evidence, and leaderboard rules as of August 2026.

Match ML and LLM interview practice to the role's artifact, the recruiter's process, and the first failed contract in a real attempt, not a fixed topic list.