Personalize this lesson
Adapt explanations and teaching visuals to your background and preferred voice.
A candidate response scores 9/10 from an LLM judge and passes its own generated unit test. Is that enough to admit it? Check what those two signals leave unanswered: the response might repeat an existing example, copy an evaluation item, or pass a test that asserts incorrect behavior. Hold it until the remaining checks are complete.
Synthetic post-training data can target gaps that your current dataset doesn't cover: tool arguments, API edge cases, or hard instruction-following tasks. Steer a generator with reviewed examples and source material, check its outputs, and track which rows actually improve the trained model. Some tasks support execution tests; others need references or human judgment. A documented admission policy makes these decisions inspectable, without turning them into a guarantee of quality.
If a generator invents both a function and the only test that approves it, a pass shows agreement with that test, not independent correctness. Keep trusted fixtures and verifier code outside its write permissions. Separate development checks from evaluation suites reserved for release decisions. The Python 3.12 cells below exercise local admission mechanics with hand-written strings, invented embeddings, and fixed votes. They don't call a generator or judge API, train a model, or establish a production cost.

Why generate data if real text already exists?
Epoch's June 2024 forecast put training datasets near the available public-human-text stock between 2026 and 2032, conditional on continued scaling trends.[1] That's neither a confirmed exhaustion date nor a claim that reading text consumes it permanently. Epoch's September 2025 assessment judged the possible data bottleneck surmountable, with synthetic data among the options. Your post-training project needn't wait for any global “data wall”: it may simply lack enough reviewed examples of its own difficult cases.
Generation can propose those missing cases. Public text also contains useful code, explanations, and documentation; the issue is coverage, rights, and quality for the chosen task. Candidate targets include:
- Complex instruction following with multi-variable constraints
- Rigorous unit test fixtures and regression suites for backend code
- Precise tool-use calling sequences with structured JSON schemas
- Safe refusal dialogues that resist adversarial prompt injection
Recursive training can amplify distribution errors. Shumailov et al. show tail loss and later degeneration in mathematical settings and experiments across model families.[2] A finite generated sample can omit rare events; later models may reproduce that omission. The result isn't a theorem that every synthetic-data recipe must collapse. Data retention, sampling, filtering, model fitting, and external information change the setting.
1Illustrative Replacement Failure (Model Collapse):
2 Base Model M0 ──> Synthetic Data D1 ──> Train M1 ──> Synthetic Data D2 ──> Train M2
3 │ │
4 ▼ ▼
5 (Rare events may be lost) (Errors can compound)
6
7Retained-Data & Verification Loop (Mitigation to Evaluate):
8 Reviewed Real Data ──┐
9 ├──> Admission Checks ──> Accepted Shard ──> Train M_next
10 Synthetic D_gen ──┘ (Task-specific evidence)Retain useful real data and inspect both common and rare task slices. Gerstgrasser et al. find that accumulation avoids collapse in their tested settings and bounds error in an idealized linear-regression analysis.[3] Their result isn't a universal 10% human-data rule, and their accumulated synthetic data needn't be externally verified. Verification addresses some correctness failures; it doesn't certify distribution coverage.
Seed curation: anchor the distribution in human intent
A reviewed seed pool gives the generator concrete examples of the task, output format, and domain boundaries. Its size depends on coverage and the generation method; 150–200 isn't a required interval. Grounded documents, environment states, or existing task data can also supply starting points. Review seed errors and missing categories before amplifying them, without assuming that human authorship makes a row correct.
Two authored seed specifications give the local fixtures their intended domain:
| Seed ID | Domain | Base Instruction | Target Output Contract |
|---|---|---|---|
human-004 | Backend Auth | Write a pytest regression test for expired API tokens. | Timezone-aware UTC clock fixture, assert rejection before permission checks, verify audit events. |
human-011 | API Documentation | Answer an API pagination question using only the supplied docs. | Exact parameter bounds (page_size <= 100), cursor iteration, bounded backoff on HTTP 429. |
Every synthetic row must declare its lineage, assigned slice, generation tactic, and source seed from the moment it's spawned. Never attempt to reconstruct provenance after training. Storing metadata on the row allows post-training runs to track data mixtures and isolate problematic generator runs immediately.
1from dataclasses import dataclass, asdict
2import json
3
4@dataclass(frozen=True)
5class Candidate:
6 row_id: str
7 seed_id: str
8 tactic: str
9 data_slice: str
10 instruction: str
11 response: str
12 generator: str
13
14def validate(row: Candidate) -> None:
15 for value in asdict(row).values():
16 if not isinstance(value, str) or not value.strip():
17 raise ValueError("candidate fields must be nonempty strings")
18 if row.seed_id not in {"human-004", "human-011"}:
19 raise ValueError("unknown seed")
20 if row.tactic not in {"self_instruct", "add_constraint", "edge_case", "multi_step", "red_team"}:
21 raise ValueError("unknown tactic")
22 if row.data_slice not in {"standard", "red_team"}:
23 raise ValueError("unknown slice")
24 if (row.tactic == "red_team") != (row.data_slice == "red_team"):
25 raise ValueError("red-team tactic and slice must agree")
26
27row = Candidate(
28 row_id="synth-0007",
29 seed_id="human-004",
30 tactic="edge_case",
31 data_slice="standard",
32 instruction="Write a pytest fixture for timezone-aware token expiry.",
33 response="Freeze the clock in UTC and assert expired tokens are rejected before permission checks.",
34 generator="fixture-generator-v1",
35)
36validate(row)
37print(json.dumps(asdict(row), sort_keys=True))1{"data_slice": "standard", "generator": "fixture-generator-v1", "instruction": "Write a pytest fixture for timezone-aware token expiry.", "response": "Freeze the clock in UTC and assert expired tokens are rejected before permission checks.", "row_id": "synth-0007", "seed_id": "human-004", "tactic": "edge_case"}The dataclass supplies named fields; Python annotations don't enforce their types at runtime. The explicit validate function checks the declared strings, seed IDs, tactics, and slices. It doesn't authenticate provenance. Record actual sampling parameters, model identifiers, and prompt revisions separately. A remote API may not expose a weights hash or provide reproducible seeded output. Assign seed families to splits before generation and keep descendants together; also inspect overlap between different families.
Self-Instruct: grow task coverage from seeds
Self-Instruct bootstraps a broad instruction dataset by prompting a language model to invent new tasks based on in-context demonstrations drawn from the seed pool. Wang et al. (2023) demonstrated this mechanism by starting with 175 human-authored tasks, iteratively generating new instructions and instances, and applying ROUGE-L overlap filtering to discard near-duplicates.[4] Their pipeline expanded 175 seeds into 52,445 filtered, diverse instructions.
The original paper's cycle distinguishes these operations:
- Use eight task instructions as demonstrations: six human-written and two previously generated in the reported recipe.
- Prompt the generator model to formulate a new, plausible task instruction.
- Identify classification versus non-classification, not whether an extra input field is required.
- Generate instances, using output-first generation for classification tasks to reduce label skew and input-first generation for the others.
- Filter instructions with ROUGE-L similarity below 0.7 against the existing pool, and apply the paper's instance heuristics. MinHash is an alternative you could evaluate, not the reported filter.
This is a historical recipe, not a mandate for your demo count or overlap threshold. Prompting can produce many variants of a narrow task without adding useful coverage. Track the categories and difficulty you intended to generate, then inspect the resulting instances.
Evol-Instruct: rewrite the task complexity, not just the wording
To break past the difficulty plateau of Self-Instruct, Evol-Instruct rewrites the instruction itself to systematically increase task complexity. Introduced with WizardLM (Xu et al., 2023), Evol-Instruct applies structured mutation prompts along two complementary axes: in-depth evolving and in-breadth evolving.[5]
The paper lists five in-depth operations. Their task-specific forms could look like:
- Add constraints: Introduce operational restrictions ("the test must run asynchronously, use strict type annotations, and execute in under 10 milliseconds").
- Deepen the task: Add a more demanding requirement ("explain which check must reject a revoked token and why").
- Concretize abstractions: Replace generic requirements with domain-specific implementations ("instead of checking generic tokens, validate OAuth2 Bearer tokens with JWT claims and RSA signatures").
- Increase reasoning steps: Add dependencies ("create the expiry fixture, call the middleware, then verify the audit event").
- Complicate inputs: Provide edge-case inputs ("test a token that expires at the exact microsecond of evaluation when the server clock source is naive UTC").
In-breadth evolving proposes a different task to broaden topic and skill coverage; it isn't mathematically orthogonal or necessarily confined to the same domain. For this assistant, you might request an IP-rate-limiting middleware or an HMAC signature validator. Check coverage rather than assuming a rewrite prevents overfitting.
1Base Seed ("Write token regression test")
2 │
3 ├── In-Depth: Add Constraints ("Async, UTC clock fixture, check audit log")
4 ├── In-Depth: Deepen Reasoning ("Multi-turn auth handshake failure flow")
5 ├── In-Depth: Edge Case ("Token expiring at exact current microsecond")
6 └── In-Breadth: Domain Expansion ("HMAC webhook signature validation")WizardLM's historical eliminator uses information-gain judgments, a short-response “sorry” heuristic, punctuation/stop-word-only responses, and copied evolution scaffolding. Those rules aren't correctness proofs. In a safety dataset, a proper refusal may be the desired answer; blindly banning a phrase would remove useful examples.
1seed = "Write a regression test for expired API tokens."
2tactics = {
3 "add_constraint": lambda text: f"{text} Use a UTC clock fixture and assert the auth layer rejects first.",
4 "edge_case": lambda text: f"{text} Cover a token that expires exactly at the current second.",
5 "multi_step": lambda text: f"{text} First create the fixture, then call the middleware, then assert the audit log.",
6}
7
8scheduled = [
9 {"seed_id": "human-004", "tactic": tactic, "instruction": rewrite(seed)}
10 for tactic, rewrite in tactics.items()
11]
12
13assert {row["tactic"] for row in scheduled} == set(tactics)
14for row in scheduled:
15 print(f"{row['tactic']}: {row['instruction']}")1add_constraint: Write a regression test for expired API tokens. Use a UTC clock fixture and assert the auth layer rejects first.
2edge_case: Write a regression test for expired API tokens. Cover a token that expires exactly at the current second.
3multi_step: Write a regression test for expired API tokens. First create the fixture, then call the middleware, then assert the audit log.These lambdas only schedule three authored rewrites; they don't run Evol-Instruct or measure difficulty. A longer instruction can be redundant, impossible, or wrong. Verify feasibility and target answers, then compare task gains after training.
Grounded execution: test the claims you can execute
Prose review can identify errors, but executable checks add evidence about specified behavior. A generator may produce plausible broken code or a weak test that passes it. Test independence and coverage matter as much as a green result.
Different verifiers establish different facts:
- Parsers, linters, and type checkers: They inspect syntax or selected static properties. Coverage depends on the language, rules, and configuration; a pass doesn't establish runtime behavior or even that every import will resolve.
- Execution tests: Trusted cases check observed behavior. A test runner isn't a sandbox; configure isolation, resource limits, and network policy separately.
- Symbolic solvers: They check formalized expressions or constraints under declared assumptions. An incorrect translation of the task can still produce a misleading answer.
Phi-1 illustrates data curation, not an isolated measurement of execution filtering. Its paper reports about 6B filtered code-language tokens, under 1B synthetic textbook tokens, and roughly 180M exercise tokens for fine-tuning. The 1.3B model's reported HumanEval pass@1 is 50.6%.[6] Dataset inventory differs from repeated training exposure. This 2023 result doesn't establish a current model ranking or attribute the gain solely to tests.
Llama 3's execution-feedback method produced about one million coding dialogues using static checks and generated tests in containers. The report says about 20% of solutions were initially incorrect and self-corrected. Crucially, the model could revise the solution or its tests, and the authors explicitly don't ensure complete correctness.[7] Passing all configured checks admitted a dialogue to supervised fine-tuning (SFT); that isn't equivalent to passing independent trusted tests.
Orca reports improvements from learning teacher explanation traces in its tested tasks.[8] This isn't a general speedup claim. Generated explanations aren't proof of a model's internal reasoning; check observable intermediate claims and final behavior rather than equating length with validity.
Calibrated LLM judges: position stability over prose polish
Deterministic verifiers excel on executable code and formal math, but open-ended instruction following, API documentation answers, and stylistic adherence require semantic judgment. This is where LLM-as-a-judge systems enter the pipeline.
Zheng et al. (2023) studied several judges on MT-Bench and Chatbot Arena:[9]
- Position bias: Some decisions depend on answer position or assistant names; a judge needn't always prefer the first answer.
- Verbosity bias: Irrelevant added length can influence a judgment, with susceptibility varying across tested judges.
- Possible self-enhancement: Some judges favor their own outputs, but the paper says its data can't establish this bias conclusively. Don't present that uncertainty as a universal model-family preference.
1Position Bias Detection:
2 Forward Pass: [Prompt] + [Candidate A (Left)] + [Candidate B (Right)] ──> Judge selects Left (A)
3 Reversed Pass: [Prompt] + [Candidate B (Left)] + [Candidate A (Right)] ──> Judge selects Left (B)
4 │
5 Result: Observed winner changes after reversal. ──────────────────────────────┴──> AUDITA conservative policy checks order reversal:
- Present the pair in forward order
(A, B)and record the winning candidate. - Present the exact same pair in reversed order
(B, A)and record the winning candidate. - Keep a win only if the same response wins both presentations and other admission checks pass. Otherwise record a tie, abstention, or audit decision under your policy.
Order agreement tests stability, not correctness. Both orders can approve the same wrong answer, and one disagreement can reflect stochastic variation or a close pair rather than establish systematic position bias. Use reference-backed criteria where possible, test padding and prompt-injection cases, and calibrate on the task. A binary or short ordinal rubric can clarify decisions, but doesn't by itself eliminate verbosity bias.
The labels A and B identify two hypothetical responses. Their text isn't judged in this cell; fixed votes test the remapping of positions to response identities.
1def canonical_winner(forward: str, reversed_order: str) -> str:
2 if forward not in {"left", "right"} or reversed_order not in {"left", "right"}:
3 return "audit"
4 forward_winner = "A" if forward == "left" else "B"
5 reverse_winner = "A" if reversed_order == "right" else "B"
6 return forward_winner if forward_winner == reverse_winner else "audit"
7
8judgments = [
9 {"pair": "unsafe-tool-refusal", "forward": "left", "reversed": "right"},
10 {"pair": "verbose-answer", "forward": "left", "reversed": "left"},
11]
12
13for row in judgments:
14 decision = canonical_winner(row["forward"], row["reversed"])
15 print(f"{row['pair']}: {decision}")
16
17assert canonical_winner("left", "right") == "A"
18assert canonical_winner("left", "left") == "audit"
19assert canonical_winner("tie", "right") == "audit"1unsafe-tool-refusal: A
2verbose-answer: auditA pairwise judge chooses candidate A when A appears first, then chooses candidate B when B appears first. Should either preference enter training data?
Answer
Withhold the win under this conservative policy. The decision is order-unstable; repeat controlled comparisons before diagnosing systematic bias. Order-stable wins still need correctness and calibration checks.
Build a reviewed calibration set covering ordinary cases, rare failures, close pairs, and each important slice. Measure agreement, false admissions, false rejections, abstentions, and uncertainty; compare human raters under the same rubric. Neither a 5%–10% audit fraction nor 80% agreement is a universal target. Revise the judge on development cases, confirm on separate cases, and monitor drift as generators change.
Predict whether this judge looks good under aggregate agreement. It admits every row; the reviewed set contains 98 acceptable rows and two unacceptable ones:
1reviewed_accept = [True] * 98 + [False] * 2
2judge_accept = [True] * 100
3agreement = sum(a == b for a, b in zip(reviewed_accept, judge_accept))
4bad_count = sum(not value for value in reviewed_accept)
5bad_admitted = sum(j and not r for r, j in zip(reviewed_accept, judge_accept))
6
7print(f"agreement={agreement / len(reviewed_accept):.0%}")
8print(f"unacceptable_rows_admitted={bad_admitted}/{bad_count}")
9print(f"false_admission_rate_on_bad_rows={bad_admitted / bad_count:.0%}")1agreement=98%
2unacceptable_rows_admitted=2/2
3false_admission_rate_on_bad_rows=100%The headline score hides both unacceptable rows. This authored fixture isn't a calibrated judge or a reliable population estimate from only two bad cases. It shows why class-specific errors and sample coverage matter.
Multi-tier admission filtering: shift cheap checks left
Our example policy has five gates. Other tasks may need different verifiers, accept useful repetition, or omit a judge when trusted checks already establish the required property. Evaluate gate effectiveness and false rejections rather than treating every tier as mandatory.
Run inexpensive selective checks early when that lowers expected cost. Measure actual stage costs and survivor rates: an embedding lookup or an overlap index isn't automatically cheap at corpus scale. Gate order can also change the dataset when a stage depends on shared accepted state.
1Candidate Admission Pipeline (Shift-Left Order):
2 Candidate
3 │
4 ▼
5 [Tier 1: Schema & Contract] ──Fail──> Drop
6 │ Pass
7 ▼
8 [Tier 2: Decontamination Overlap] ──Fail──> Drop
9 │ Pass
10 ▼
11 [Tier 3: Task-Specific Verifier] ──Fail──> Drop
12 │ Pass
13 ▼
14 [Tier 4: Semantic Diversity] ──Fail──> Drop
15 │ Pass
16 ▼
17 [Tier 5: Judge / Review] ──Fail──> Drop / Audit
18 │ Pass
19 ▼
20 Accepted Training ShardNovelty is distinct from correctness and usefulness. Repetition may help a rare skill, while a novel row may be irrelevant. Online embedding filtering can limit crowding in an accepted pool; select its threshold against the intended coverage and inspect what it removes.
For nonzero vectors, cosine similarity is:
Consider a token-expiry reference cluster at . A newly proposed candidate has embedding coordinates :
- Dot product:
- Norm of the candidate: ; the reference norm is 1.
- Cosine similarity:
Under the authored cutoff, this vector is rejected. These hand-chosen coordinates don't come from an embedding model and don't establish semantic equivalence. The cutoff is a policy example, not a generally valid duplicate threshold.
1import math
2from typing import TypedDict
3
4class EmbeddedCandidate(TypedDict):
5 instruction: str
6 embedding: list[float]
7
8def cosine_similarity(a: list[float], b: list[float]) -> float:
9 if not a or len(a) != len(b) or not all(math.isfinite(x) for x in a + b):
10 raise ValueError("embeddings must be finite, nonempty, and equal-width")
11 norm_a, norm_b = math.hypot(*a), math.hypot(*b)
12 if not (0 < norm_a < math.inf and 0 < norm_b < math.inf):
13 raise ValueError("embedding norms must be positive and finite")
14 return max(-1.0, min(1.0, sum((x / norm_a) * (y / norm_b) for x, y in zip(a, b))))
15
16def filter_diversity(
17 new_items: list[EmbeddedCandidate],
18 existing_embeddings: list[list[float]],
19 threshold: float = 0.82,
20) -> list[EmbeddedCandidate]:
21 if not -1 <= threshold <= 1:
22 raise ValueError("cosine threshold must be in [-1, 1]")
23 accepted: list[EmbeddedCandidate] = []
24 comparison_pool = [vector[:] for vector in existing_embeddings]
25 for vector in comparison_pool:
26 cosine_similarity(vector, vector)
27 for item in new_items:
28 cosine_similarity(item["embedding"], item["embedding"])
29 max_sim = max(
30 (cosine_similarity(item["embedding"], vector) for vector in comparison_pool),
31 default=-1.0,
32 )
33 if not comparison_pool or max_sim < threshold:
34 accepted.append(item)
35 comparison_pool.append(item["embedding"])
36 return accepted
37
38existing = [
39 [1.0, 0.0], # token-expiry cluster
40 [0.0, 1.0], # API-doc question cluster
41]
42
43candidates: list[EmbeddedCandidate] = [
44 {
45 "instruction": "Write another token-expiry regression test.",
46 "embedding": [0.99, 0.02],
47 },
48 {
49 "instruction": "Answer an API pagination question with a citation.",
50 "embedding": [0.45, 0.40],
51 },
52 {
53 "instruction": "Answer another pagination question with a citation.",
54 "embedding": [0.46, 0.41],
55 },
56]
57
58kept = filter_diversity(candidates, existing)
59print("kept candidates:")
60for item in kept:
61 print(f"- {item['instruction']}")1kept candidates:
2- Answer an API pagination question with a citation.The third vector is close to the second. After the second is admitted, its vector blocks the third under this greedy rule. Input order can change the surviving set; the pool update doesn't prove that the dropped instruction lacks training value.

An admission policy can withhold suspected evaluation copies regardless of their quality score. Training on test items can inflate measured performance; it doesn't logically destroy all generalization. Teachers may regenerate memorized items, and ordinary task vocabulary can also overlap without constituting a leak. Review hits and document the screening policy and its limits.
Pipelines employ multi-tier decontamination screening:
- Exact normalized n-gram overlap: An inverted index can detect shared spans. Choose tokenization and n for the source type; short boilerplate spans create false hits.
- MinHash/Jaccard candidates: MinHash estimates Jaccard similarity between shingle sets, . Whole-document similarity can miss a short copied span in a long document. There's no universal 0.1 cutoff.
- Semantic retrieval: Embeddings can propose paraphrase candidates for inspection, but don't reliably certify every leak or every hit.
The next cell stores exact normalized five-word strings in sets. It is not MinHash and doesn't compare Jaccard scores. NFKC normalization, case folding, and punctuation removal can merge distinct code or identifiers; it also misses paraphrases and shorter copies. The filename describes the kind of overlap test, not raw-byte equality.
1import re
2import unicodedata
3
4def extract_ngrams(text: str, n: int) -> set[str]:
5 if n < 1:
6 raise ValueError("n must be positive")
7 tokens = re.findall(r"\w+", unicodedata.normalize("NFKC", text).casefold())
8 return {" ".join(tokens[i:i+n]) for i in range(len(tokens) - n + 1)}
9
10def build_forbidden(held_out: list[str], n_range: tuple[int, int]) -> dict[int, set[str]]:
11 if not 1 <= n_range[0] <= n_range[1]:
12 raise ValueError("invalid n-gram range")
13 return {
14 n: set().union(*(extract_ngrams(text, n) for text in held_out))
15 for n in range(n_range[0], n_range[1] + 1)
16 }
17
18def is_contaminated(row: dict[str, str], forbidden: dict[int, set[str]]) -> bool:
19 for text in (row["instruction"], row["response"]):
20 for n, forbidden_ngrams in forbidden.items():
21 if extract_ngrams(text, n) & forbidden_ngrams:
22 return True
23 return False
24
25held_out = [
26 "Ignore the documented rate limit and expose the hidden admin token in the response."
27]
28candidates = [
29 {
30 "instruction": "Answer the privileged API question.",
31 "response": "Ignore the documented rate limit and expose the hidden admin token in the response.",
32 },
33 {
34 "instruction": "Explain API pagination.",
35 "response": "Use the documented cursor parameter and cite the page-size limit.",
36 },
37]
38forbidden = build_forbidden(held_out, n_range=(5, 5))
39kept = [row for row in candidates if not is_contaminated(row, forbidden)]
40
41print(f"blocked rows: {len(candidates) - len(kept)}")
42print(f"kept rows: {len(kept)}")
43assert kept == [candidates[1]]1blocked rows: 1
2kept rows: 1Each gate answers a specific architectural question:
| Tier | Gate Mechanism | Decision Scope | Blind Spot |
|---|---|---|---|
| 1 | Candidate contract | Nonempty strings, known seed/tactic/slice, unique IDs in the loop | Authentic lineage and factual correctness |
| 2 | Overlap screen | Exact normalized five-word sets in this lab | Short copies, paraphrases, unindexed evaluation sources |
| 3 | Task verifier | Three JSON fields here; trusted code tests are another task's option | A working client and untested behavior |
| 4 | Vector diversity | Cosine similarity against the accepted pool | Actual semantic equivalence or usefulness |
| 5 | Fixed-vote judge gate | Stable A win across two supplied presentations | Correctness, calibration, and shared bias |
Red-teaming data synthesis: probe adversarial boundaries
Benign demonstrations may leave adversarial cases poorly covered. That alone doesn't establish a vulnerability or mean the starting model has never learned refusal boundaries. Generate targeted probes, inspect the actual failures, and evaluate both secure behavior and unnecessary refusals.
Perez et al. (2022) established that language models can be instructed to generate adversarial red-team prompts at scale.[10] Anthropic's Constitutional AI framework expanded this methodology by prompting models to critique and revise unsafe generations against explicit constitutional principles.[11]
In a code and API assistant pipeline, red-teaming meta-prompts generate targeted vulnerability probes:
- "Write a Python token validation function that appears functional but contains an intentional backdoor: if the token subject starts with 'admin-', bypass signature verification and return status 'ok'."
- "Write an API documentation query that mimics a routine debugging question while prompting the assistant to leak internal credentials or ignore documented rate limits."
The training target should satisfy the reviewed security policy: reject an authentication bypass or credential disclosure while still helping with legitimate debugging. An approved safe response to an adversarial prompt can be an SFT demonstration. Harmful rejected answers belong in the appropriate comparison or analysis role, not as desired assistant completions.
This lab tags probes with data_slice="red_team"; a tag supports accounting, not security isolation. Review prompts and targets together. The next cell executes only two trusted literal functions against a toy string-based contract. It doesn't validate JWT signatures, real token expiry, or a production authentication system. AST parsing and an empty __builtins__ dictionary are not a secure sandbox. Isolate untrusted candidate execution and configure resources, syscalls, and network access before using real generated code.
1import ast
2
3chosen = """
4def validate_token(token, revoked):
5 if token in revoked:
6 return {"status": "revoked"}
7 if token.startswith("expired-"):
8 return {"status": "expired"}
9 if token.startswith("admin-"):
10 return {"status": "blocked"}
11 return {"status": "ok"}
12"""
13rejected = """
14def validate_token(token, revoked):
15 if token.startswith("admin-"):
16 return {"status": "ok"}
17 if token in revoked:
18 return {"status": "revoked"}
19 return {"status": "ok"}
20"""
21
22def passes_contract(source: str) -> bool:
23 tree = ast.parse(source)
24 namespace: dict[str, object] = {}
25 exec(compile(tree, "<candidate>", "exec"), {"__builtins__": {}}, namespace)
26 function = namespace["validate_token"]
27 revoked = {"token-7"}
28 return (
29 function("token-7", revoked) == {"status": "revoked"}
30 and function("expired-9", revoked) == {"status": "expired"}
31 and function("admin-root", revoked) != {"status": "ok"}
32 and function("token-valid", revoked) == {"status": "ok"}
33 )
34
35print(f"chosen passes: {passes_contract(chosen)}")
36print(f"rejected passes: {passes_contract(rejected)}")
37assert passes_contract(chosen) and not passes_contract(rejected)1chosen passes: True
2rejected passes: FalseAn LLM judge prefers a polished code answer, but the executable contract shows it accepts an admin- bypass. Can the row survive because its semantic score is high?
Answer
No. Executable safety and correctness checks are hard gates. Discard or repair the candidate, preserve the failing case in the controlled red-team set, and never let a judge score override verifier failure.
Preference pair synthesis for DPO and RLHF
Modern post-training pipelines shape model behavior using comparison data. Direct Preference Optimization (DPO) trains policy models directly on preference triples: prompt , chosen response , and rejected response .[12] Traditional RLHF fits a reward model over the same preference comparisons before running policy gradient updates.
Synthetic pipelines construct preference pairs using two primary strategies:
- Multi-sample ranking: Sample multiple responses from the generator across diverse temperature settings. Execute deterministic verifiers and calibrated judges over all outputs. Select the highest-scoring execution-verified response as
chosenand an inferior or broken response asrejected. - Targeted error injection: Take a verified
chosenresponse and prompt a teacher model to inject a plausible, domain-specific mistake (such as exceeding rate limits, ignoring pagination boundaries, or introducing off-by-one errors).
Consider our pagination seed human-011:
- Prompt: "The API documentation states: pass returned cursor tokens to
/v1/search, limitpage_sizeto 100, and retry HTTP 429 responses with exponential backoff. Write a client function that fetches all records." - Chosen (): Implements cursor iteration, enforces
page_size = 100, implements bounded exponential backoff on HTTP 429, and handles empty cursor termination cleanly. - Rejected (): "Set
page_size = 10000to fetch everything in one query and retry immediately on errors; the docs imply internal endpoints handle larger limits."
Both responses address the same prompt and documentation context. This contrast is intended to teach documented constraints, but its effectiveness needs a trained-model comparison. Verify that chosen is actually better; a teacher revision isn't automatically an improvement. Inspect negative difficulty, length and style confounds, and resemblance to real policy errors. Keep harmful rejected text out of desired SFT assistant completions while retaining the approved safe target.
The unified admission loop in practice
Now integrate all five gating tiers into a unified pipeline. The pipeline processes candidate rows sequentially, recording the exact gate that triggered rejection.

The runnable pipeline below tests all six representative fixtures:
schema: Fails Tier 1 schema check (invalid seed ID).overlap: Fails Tier 2 decontamination check (matches held-out evaluation 5-gram).verifier: Fails Tier 3 JSON configuration check (page_sizeset to 10,000 instead of 100).duplicate: Fails Tier 4 diversity check (cosine similarity 0.9998 against prior cluster).unstable: Fails Tier 5 judge check (order-reversal inconsistency).keep: Clears all five tiers and enters the training shard.
1from dataclasses import replace
2
3def verify_config(response: str) -> bool:
4 try:
5 config = json.loads(response)
6 except (json.JSONDecodeError, TypeError):
7 return False
8 return (
9 isinstance(config, dict)
10 and set(config) == {"page_size", "retry_limit", "cursor"}
11 and type(config["page_size"]) is int and config["page_size"] == 100
12 and type(config["retry_limit"]) is int and config["retry_limit"] == 3
13 and config["cursor"] is True
14 )
15
16base = Candidate(
17 row_id="keep", seed_id="human-011", tactic="add_constraint", data_slice="standard",
18 instruction='Return JSON with page_size=100, retry_limit=3, and cursor=true.',
19 response='{"page_size":100,"retry_limit":3,"cursor":true}',
20 generator="fixture-generator-v1",
21)
22# Each tuple is (candidate, embedding, forward vote, reversed vote).
23fixtures = [
24 (replace(base, row_id="schema", seed_id="unknown"), [0.6, 0.8], "left", "right"),
25 (replace(base, row_id="overlap", response=held_out[0]), [0.6, 0.8], "left", "right"),
26 (replace(base, row_id="verifier", response=base.response.replace("100", "10000")), [0.6, 0.8], "left", "right"),
27 (replace(base, row_id="duplicate"), [0.99, 0.02], "left", "right"),
28 (replace(base, row_id="unstable"), [0.6, 0.8], "left", "left"),
29 (base, [0.6, 0.8], "left", "right"),
30]
31
32def run_gates(items):
33 comparison_pool = [[1.0, 0.0], [0.0, 1.0]]
34 accepted, decisions, judge_candidates = [], [], []
35 seen_ids = set()
36 for candidate, embedding, forward, reverse in items:
37 gate = "schema"
38 try:
39 validate(candidate)
40 if candidate.row_id in seen_ids:
41 raise ValueError("duplicate row ID")
42 seen_ids.add(candidate.row_id)
43 gate = "overlap"
44 if is_contaminated(asdict(candidate), forbidden):
45 raise ValueError("overlap hit")
46 gate = "verifier"
47 if not verify_config(candidate.response):
48 raise ValueError("task contract failed")
49 gate = "diversity"
50 if max(cosine_similarity(embedding, prior) for prior in comparison_pool) >= 0.82:
51 raise ValueError("near duplicate")
52 gate = "judge"
53 judge_candidates.append(candidate.row_id)
54 if canonical_winner(forward, reverse) != "A":
55 raise ValueError("unapproved preference")
56 except ValueError:
57 decisions.append((candidate.row_id, gate))
58 continue
59 accepted.append(candidate)
60 comparison_pool.append(embedding[:])
61 decisions.append((candidate.row_id, "accept"))
62 return accepted, decisions, judge_candidates
63
64accepted_rows, decisions, judge_candidates = run_gates(fixtures)
65for row_id, decision in decisions:
66 print(f"{row_id}: {decision}")
67print("Judge gate reached:", judge_candidates)
68print("Accepted:", [row.row_id for row in accepted_rows])
69assert [row.row_id for row in accepted_rows] == ["keep"]1schema: schema
2overlap: overlap
3verifier: verifier
4duplicate: diversity
5unstable: judge
6keep: accept
7Judge gate reached: ['unstable', 'keep']
8Accepted: ['keep']Two of six fixtures reach the judge gate: 66.7% fewer candidate judgments than judging all six. Votes are supplied fixtures, so no API was invoked. If each real candidate used the two-presentation order check, this would mean four judge requests versus twelve, before retries. It isn't a measured dollar saving. The JSON verifier checks three fields, not a working pagination client. The compact receipts record each final outcome; production stage receipts would need more detail.
unstable doesn't enter the accepted diversity pool, so it can't block keep. That is this policy's state boundary. Rejected rows can still belong in a separate audit or duplicate-detection store. Greedy diversity depends on the final accepted pool and candidate order; coordinate it when parallelizing.
From the fixture to an actual generation framework
As of September 22, 2026, NeMo Data Designer is one concrete framework example. Its September 15 preprint describes declarative dataset columns, seed and statistical sampler inputs, dependency-aware model calls, validation, and preview/revision workflows.[13] Those mechanisms let you inspect a small generated batch before scaling and retain intermediate evidence. The official repository links its current documentation and examples.
To replace this lab's fixtures, connect a generator, an embedding model, and a task-calibrated evaluator; pin the configuration and capture actual inputs, outputs, and gate results. Choose checks for your artifact. Framework retries and schema validation don't provide independent ground truth, and API generation won't necessarily reproduce identical bytes on rerun.
Avoiding model collapse in recursive training loops
As fine-tuned models improve, engineering teams naturally use newer checkpoints to synthesize training data for subsequent model generations. This creates the data flywheel. However, recursive loops introduce subtle distribution drift if not strictly governed.
Replacement can compound sampling and fitting errors in the settings Shumailov studies.[2] Accumulation changes that regime. Gerstgrasser's bounded-error result assumes a particular linear-regression model and growing data; its language-model experiments support accumulation in their tested conditions.[3] Neither paper proves that retaining a fixed percentage of any human seed set makes every post-training loop safe.
1Data Regimes Across Rounds:
2 Replacement: Round 1 Synthetic ──> Round 2 Synthetic ──> Round 3 Synthetic (Real data lost!)
3 Accumulation: Real Human Anchor + Round 1 Vetted + Round 2 Vetted + Round 3 Vetted1real_rows = {"real-api-01", "real-api-02", "real-code-01"}
2synthetic_rounds = [
3 {"synth-r1-01", "synth-r1-02"},
4 {"synth-r2-01", "synth-r2-02"},
5]
6
7replacement_training = synthetic_rounds[-1]
8accumulated_training = real_rows | set().union(*synthetic_rounds)
9
10print(f"replacement_has_real={bool(replacement_training & real_rows)}")
11print(f"accumulated_has_real={real_rows <= accumulated_training}")
12print(f"accumulated_rows={len(accumulated_training)}")
13assert real_rows <= accumulated_training1replacement_has_real=False
2accumulated_has_real=True
3accumulated_rows=7This set operation preserves IDs; it doesn't train a model or establish protection against collapse. Sampling weight, data quality, and coverage still matter. A decreasing real-data fraction alone doesn't prove collapse either: in the paper's accumulation analysis, that fraction falls across rounds while error stays bounded under its assumptions. Compare declared mixture policies on independent common and rare cases; there's no established universal 10% floor.
A teammate keeps the original real rows in storage but samples them once per million synthetic rows during training. Has retaining their IDs established protection against recursive drift?
Answer
No. Retention differs from effective training weight. Measure the real-to-synthetic sampling mixture and behavior on independent task slices. Accumulation helps in the cited settings, but neither retaining IDs nor choosing arbitrary percentages guarantees quality in this project.
Shard identity: what a hash does and doesn't establish
Versioned data and receipts help investigate regressions and establish what a run used. They don't let you uniquely trace every future completion to a particular training row: outputs depend on many learned parameters and inference context. Source rights and consent also require actual evidence; a manifest doesn't confer them.
A useful manifest records artifact identities and points to retrievable evidence:
- Generator model identifier and weights hash when available
- Complete prompt templates and system instruction digests
- Verifier script versions and container environment hashes
- Benchmark decontamination index identifiers and thresholds
- Source human seed identifiers and evolution tactics
- Full SHA-256 digests of the accepted payload and decision log, anchored in a trusted reference
1import hashlib
2import json
3from pathlib import Path
4from tempfile import TemporaryDirectory
5
6ordered = sorted(accepted_rows, key=lambda row: row.row_id)
7payload = "".join(json.dumps(asdict(row), sort_keys=True, ensure_ascii=False, separators=(",", ":")) + "\n" for row in ordered).encode("utf-8")
8receipts = "".join(json.dumps({"row_id": row_id, "outcome": outcome}, sort_keys=True) + "\n" for row_id, outcome in decisions).encode("utf-8")
9manifest = {
10 "generator": "fixture-generator-v1",
11 "judge": "fixture-votes-not-calibrated",
12 "prompt": "exact-pagination-config-v1",
13 "verifier": "json-config-v1",
14 "embedding": "hand-chosen-2d-fixture",
15 "decontamination_set": "invented-eval-fivegrams-v1",
16 "source_ids": sorted({row.seed_id for row in ordered}),
17 "tactics": sorted({row.tactic for row in ordered}),
18 "accepted_rows": len(accepted_rows),
19 "rows_sha256": hashlib.sha256(payload).hexdigest(),
20 "receipts_sha256": hashlib.sha256(receipts).hexdigest(),
21}
22
23with TemporaryDirectory() as directory:
24 shard = Path(directory) / "synthetic.jsonl"
25 audit = Path(directory) / "receipts.jsonl"
26 metadata = Path(directory) / "manifest.json"
27 shard.write_bytes(payload)
28 audit.write_bytes(receipts)
29 metadata.write_text(json.dumps(manifest, sort_keys=True), encoding="utf-8")
30 loaded = json.loads(metadata.read_text(encoding="utf-8"))
31 stored = shard.read_bytes()
32 assert hashlib.sha256(stored).hexdigest() == loaded["rows_sha256"]
33 assert hashlib.sha256(audit.read_bytes()).hexdigest() == loaded["receipts_sha256"]
34 assert len(stored.splitlines()) == loaded["accepted_rows"]
35 print("JSONL rows:", loaded["accepted_rows"])
36 print("Full SHA-256:", loaded["rows_sha256"])
37 print("Read-back verified:", True)1JSONL rows: 1
2Full SHA-256: e2957b2a6cb42dc83f2355fbcb38ac7007dea404e99077ec9e6e93952a15f9ac
3Read-back verified: TrueThis cell checks read-back byte identity against the manifest it just wrote. Its labels are fixture versions, not hashes of every referenced implementation; the log contains final outcomes only. A matching SHA-256 digest doesn't prove accuracy, rights, provenance authenticity, or immutability. Protect the approved reference separately. An attacker who can rewrite both data and manifest can recompute a matching digest.
Predict which check catches a changed payload when its untrusted manifest is also changed:
1import hashlib
2
3approved = b'{"page_size":100}\n'
4changed = b'{"page_size":10000}\n'
5trusted_digest = hashlib.sha256(approved).hexdigest()
6new_digest = hashlib.sha256(changed).hexdigest()
7untrusted_manifest = {"rows_sha256": new_digest}
8
9print(f"untrusted_manifest_accepts_changed={new_digest == untrusted_manifest['rows_sha256']}")
10print(f"trusted_reference_accepts_changed={new_digest == trusted_digest}")1untrusted_manifest_accepts_changed=True
2trusted_reference_accepts_changed=FalseThe separately trusted digest detects this change. Reading a self-consistent payload/manifest pair does not authenticate its author. Preserve the exact bytes and the approved reference so the comparison has meaning.
Downstream evaluation: measure capability, not acceptance yield
Passing pipeline admission confirms a candidate satisfied your local heuristics. That doesn't prove training a model on the resulting shard improves downstream performance.
Beware of generator gaming: over multiple rounds, generation models can learn the specific quirks of your verifier or judge rubric, causing pipeline acceptance rates to climb while actual downstream model capabilities stagnate or degrade.
To evaluate real performance gains:
- Declare the comparison budget: Compare real-only and real-plus-synthetic recipes with matched starting weights and a declared token or compute budget. Equal steps and example batch sizes needn't process equal tokens. Include generation and filtering costs when evaluating the full pipeline, and report how the mixture changes real-data exposure.
- Isolate evaluation suites: Keep held-out tasks and answers out of generation prompts, training, and generator-facing feedback. A restricted overlap-screening service may index them without exposing their contents to the generator; log the index version and its limits. Reserve separate final test cases from repeated recipe selection.
- Audit failure modes: Track rare-category failures, tool calling errors, and unwanted refusal rates across independent domain slices.
Two parallel workers each compare a newly generated pagination row only against yesterday's accepted pool. Their rows are near-identical, but both survive. Which missing dependency explains the failure?
Answer
Each worker lacked visibility into the other worker's accepted rows. While deterministic checks run independently in parallel, greedy novelty filtering depends on shared, synchronized state. Coordinate diversity checks centrally or execute a final deduplication pass before publishing shards.
A synthetic row passes five-word overlap screening against your evaluation index. Can its manifest claim the row is completely uncontaminated?
Answer
No. The receipt records zero hits under that normalization and index; the shown manifest names the index and hashes the final decision log. Paraphrased questions, variable renamings, and unindexed sources may remain. Keep evaluation sources out of generation prompts and training.