LeetLLM
My PlanLearnGlossaryTracksPracticeBlog
LeetLLM

Your go-to resource for mastering AI & LLM systems.

Product

  • Learn
  • Glossary
  • Tracks
  • Practice
  • Blog
  • RSS

Legal

  • Terms of Service
  • Privacy Policy

© 2026 LeetLLM. All rights reserved.

Blog

AI Engineering Blog

Practical notes on LLM systems, evaluation, agents, inference, and developer workflows.

19 postsRSS
LeetLLMContent EngineeringAI Engineering+2

How We Built LeetLLM

How LeetLLM keeps a 196-lesson visible curriculum coherent: Git-backed bundles, research packets, themed visual builds, bounded validators, and a no-traffic Cloud Run candidate before readers arrive.

Jun 21, 202614 min
Read
AI CodingDeveloper ToolsCost Optimization+5

Best AI Coding Plans in 2026

Compare 2026 AI coding plan prices, usage meters, model access, and review surfaces across Cursor, Claude Code, Codex, Copilot, Devin, Gemini Code Assist, and secondary model plans.

Jun 6, 202616 min
Local LLMQwen3.6Unsloth+3

Run Qwen3.6 Locally with Unsloth GGUF

Run Qwen3.6 locally with llama.cpp and Unsloth GGUF. Pin 27B or 35B-A3B, budget total memory from Unsloth's ladder, expose a local endpoint, then test MTP only after the plain run works.

May 13, 202621 min
CareerPortfolioProjects

AI Engineer Portfolio Projects for Interviews

Five AI engineering project ideas built around evidence a reviewer can inspect: eval rows, traces, tests, costs, failures, and design decisions.

May 9, 202614 min
CareerAI EngineeringRoadmap

How to Become an AI Engineer from Zero in 2026

Pick an AI engineering role, grow one incident-assistant project from tests to measured deployment, and collect evidence that matches that starting point.

May 9, 202615 min
DeepSeekOpen ModelsAI Infrastructure+2

DeepSeek V4: Facts, Claims, and Fit

Hosted DeepSeek V4, the open checkpoints, and the Hugging Face parameter badges are different artifacts. Pin the exact build before you route traffic.

Apr 27, 202612 min
OpenClawAI Coding PlansCost Optimization+4

Best AI Plans for OpenClaw in 2026

Compare ChatGPT/Codex, MiniMax, Qwen Cloud, Z.AI, and direct OpenAI API billing for OpenClaw. Prices, quotas, setup, and fallback risks verified August 2026.

Apr 4, 202617 min
Local LLMOllamaGemma 4+2

Run Gemma 4 Locally with Ollama

Run Gemma 4 locally with Ollama: pick an exact tag, leave memory for context, verify placement with ollama ps, and benchmark the native API.

Apr 2, 202615 min
InferencevLLMSGLang+5

vLLM vs SGLang vs TensorRT-LLM vs Ollama: Choosing an Inference Engine in 2026

Compare vLLM, SGLang, TensorRT-LLM, Ollama, and llama.cpp by deployment surface, checkpoint recipes, latency, throughput, and operating cost.

Apr 1, 202615 min
AI EngineeringInterview PreparationArchitecture+1

50 LLM Interview Questions for 2026

Fifty LLM engineering questions that connect model mechanics to serving, retrieval, agents, evaluation, and safety through evidence and failure diagnosis.

Mar 21, 202636 min
AI EngineeringToolsDeep Dive+1

AI Coding Assistants in 2026

Compare Cursor, Codex, Copilot, and Claude Code through one authorization change. Choose by operating mode, review surface, cost meter, and team controls.

Mar 16, 202611 min
CareerCompensation

AI Engineer Salary Guide 2026

Compare AI engineering pay without mixing BLS wages, posted salary, and total compensation. Match role family, level, location, equity terms, and source date first.

Mar 16, 202613 min
LLMsAI EngineeringDeep Dive+2

Open-Weight vs Closed API LLMs

Pick an LLM lane by data boundary and license first, then compare task quality, serving evidence, and cost only among the options you're allowed to run.

Mar 16, 202615 min
Context WindowsLong ContextBenchmarks+2

Million-Token Context Windows

Million-token windows can hold a bounded corpus, but capacity doesn't guarantee reliable retrieval. Measure effective context, watch length-based prices, and choose between full context and RAG.

Mar 14, 202612 min
AgentsDeep DiveTutorial

How to Build an AI Agent from Scratch

Build a working AI agent as a plain Python loop: the model proposes a tool, your runtime validates and executes it, and the observation drives the next turn.

Feb 19, 202614 min
AI EngineeringRAGFine-Tuning+1

RAG vs Fine-Tuning vs Prompting

Choose prompting, retrieval, fine-tuning, or a hybrid by tracing each failed eval case to missing evidence, unclear instructions, recurring behavior, or a hard application rule.

Feb 19, 202616 min
AI EngineeringCareerIndustry

What Does an AI Engineer Actually Do?

AI engineer isn't one job. Applied product work, ML platforms, and research engineering ship different artifacts. Live postings show how to tell them apart.

Feb 19, 202612 min
BenchmarksEvaluationSWE-bench+2

Understanding SWE-bench

SWE-bench scores a system, not a model. Walk through the task format, harness, variants, Verified set, contamination evidence, and leaderboard rules as of August 2026.

Feb 17, 202617 min
CareerInterview Prep2026

How to Prepare for ML & LLM Engineering Interviews in 2026

Match ML and LLM interview practice to the role's artifact, the recruiter's process, and the first failed contract in a real attempt, not a fixed topic list.

Feb 16, 202614 min