[home] [coding projects] [research projects] [research interests]
A manual for curriculum-guided local corpus generation, filtering, and evaluation.
This project is a corpus-construction pipeline, not merely a loop that asks a language model for more text. It separates planning, curriculum construction, generation, filtering, and evaluation so that each stage can be inspected or replaced independently.
The guiding idea is that a seed example contains more than wording. It has an artifact type, a domain, skills, useful transformations, an audience, a difficulty level, stylistic constraints, and a structural pattern. The pipeline tries to extract those properties first and generate from that representation rather than repeatedly paraphrasing the original text.
| 1. Pipeline 2. Planning 3. Generation modes |
4. Filtering 5. Running the generator 6. Evaluating the corpus |
The generator can use Ollama, a local Hugging Face model, an arbitrary command-line process, or a mock backend for testing. Provider choice is isolated behind the same client interface, so the planning and filtering logic does not depend on one particular model runtime.
Every seed receives a structured plan. The plan is intentionally richer than a topic label:
@dataclass
class Plan:
source_id: str
artifact_type: str
domain: str
skills: List[str]
useful_operations: List[str]
avoid_operations: List[str]
possible_audiences: List[str]
difficulty: str
style_constraints: List[str]
structure: List[str]
notes_for_generator: str
If model-produced planning fails, the modular planner has a deterministic fallback with generic skills, useful operations, audiences, and a setup/development/conclusion structure. A global curriculum is then built from the collection of seed plans. That summary supplies the “frontier” jobs later in the pipeline.
A local job chooses one seed and one operation that its plan says is useful. A recombine job samples two or three plans and asks for a new record that combines useful traits without copying their wording. A frontier job has no direct seed at all and samples from the global curriculum.
for i in range(local_n):
sid = rng.choice(seed_ids)
plan = plans_by_seed.get(sid, {})
op = rng.choice(_ops_from_plan(plan, fallback_ops))
for i in range(recombine_n):
k = min(rng.choice([2, 2, 3]), len(seed_ids))
sids = rng.sample(seed_ids, k=k)
for i in range(frontier_n):
op = rng.choice(fallback_ops)
The requested ratios are normalized to the target record count and the jobs are shuffled before generation. This avoids generating the entire corpus in three large homogeneous blocks.
A generated record must survive a common filter regardless of how it was produced. The filter checks length, banned model/meta phrases, fake-citation risk, repeated lines, repeated n-grams, similarity to seeds, and similarity to records that have already been accepted.
if wc < cfg.min_words:
reasons.append("too_short")
if wc > cfg.max_words:
reasons.append("too_long")
if _repeated_line_fraction(stripped) > cfg.max_repeated_line_fraction:
reasons.append("repeated_lines")
if _repeated_ngram_fraction(stripped) > cfg.max_repeated_ngram_fraction:
reasons.append("repeated_ngrams")
for seed in seed_texts:
sim = jaccard(
shingles,
shingle_set(seed, cfg.shingle_n)
)
if sim > cfg.max_seed_jaccard:
reasons.append(f"too_close_to_seed:{sim:.3f}")
break
for old in existing_shingles:
sim = jaccard(shingles, old)
if sim > cfg.max_existing_jaccard:
reasons.append(f"near_duplicate:{sim:.3f}")
break
This stage is important because the goal is not simply to maximize generation count. The output should move away from the seed set without collapsing into near-duplicates, boilerplate, or invented citations.
The supplied start script creates a 10,000-record dataset from a seed file through a local Ollama server and includes the seed material in the compiled output:
bash scripts/start_data.sh
The corresponding resume script uses the same configuration with --resume, so plans and accepted records already written to disk can be reused after an interrupted run.
The repository also contains a small Transformer training harness and separate scripts for control, arxiv, and synthetic corpora. This turns corpus construction into an experiment rather than a one-way export: the generated data can be compared under the same model size, sequence length, optimization settings, and validation text.
bash scripts/synthetic.sh
The training report includes train and validation loss, perplexity, bits per byte, train/validation gap, overfit ratio, loss improvement, and throughput. Those diagnostics are useful for asking whether the generated corpus changes generalization rather than merely whether the generator produced a large file.
| Path | Purpose |
|---|---|
| generator/planning.py | Seed plans and curriculum summary. |
| generator/jobs.py | Local / recombine / frontier scheduling. |
| generator/llm_clients.py | Model backend abstraction. |
| generator/filters.py | Quality, repetition, citation, and similarity rejection. |
| generator/make_synthetic_dataset.py | End-to-end generation and resume logic. |
| exp/exp_main.py | Language-model evaluation of resulting corpora. |
These pages are selective technical manuals: enough source to expose the mechanism, not a mirror of the entire repository.
Last updated: September 14, 2026.