LLM Fine-Tuning: Your Dataset Is the Hard Part, Not the Run
For an internal smart-contract audit SFT lane, hold the run until the dataset passes dedup, leakage, grounding, negative-coverage, and truncation gates.
Filed under Local LLM Inference· see every report on this topic
Verdict: For a security engineer building an internal smart-contract-audit assistant, halt before training. The base model may handle the domain badly, but GPU-hours will not repair a dataset that leaks answers, loses rows, repeats artifacts, omits negatives, truncates reasoning, or cannot locate source. Fix those gates first.
This June–July 2026 field report covers an internal Solidity repository-audit SFT pipeline. It does not show a trained model beating a baseline: there is no verified adapter artifact, evaluation win, or benchmark delta. It records failures caught before training and the checks that caught them.
The operator and the leak
A small security or developer-tools team can export ChatML and still have no usable training set. Parsing proves only that the file serializes. Audit examples must teach source-backed inspection and restraint, not label copying.
Our planned pipeline built blind prompts and grounded targets from Solidity repositories:
- Selector
- Parse Selector Output
- Blind Validation
- Target Writer
- Target Sanitization
- Citation Repair
- Target Validation
- ChatML Build + Export
([GBrain: projects/bastet-pipeline, July 2026])
The v0 output looked healthy: 6.0M training tokens from 26 repositories, 77% of records scoring 80–100, and yearn plus balancer held out for validation. The same pipeline later exposed a deduplication defect that retained only 287 records from a much larger candidate set. A score is not a dataset audit.
Run the dataset gates before the training recipe
For audit work, the prompt contains source and the target contains the report. The prompt must not contain the finding title, vulnerability class, or label ([GBrain: projects/bastet-pipeline, July 2026]).
The selector and target writer used agent_input.md, with a 120K-character inline budget and inputs of roughly 36–90KB. Shell ARG_MAX was approximately 100–200KB, so the pipeline sent prompts through stdin or temporary files rather than command-line arguments ([GBrain: projects/bastet-pipeline, July 2026]).
| Gate | Receipt | Decision |
|---|---|---|
| Dedup | 191 exact duplicates removed; later hash bug kept 287 records | Count rows, artifacts, and near-duplicate prompts separately |
| Labels | Re+AC0-erntrancy, unmapped labels, High-severity bias |
Normalize before stratifying |
| Negatives | 13 positive-only splits; later SolidiFI v1.5, 2,187 rows with 19.98% negatives | Include clean, fixed, and false-positive cases |
| Truncation | 137 train, 39 validation; 26%/33% over 4K; 32% train truncated | Measure the cutoff, not the average |
| Leakage | Finding titles, classes, or labels in prompts | Hard-block the record |
| Grounding | Web3Bugs reports could not localize to source | Require candidate recall above 70% |
| Holdout | yearn and balancer |
Hold out repositories, not random rows |
The label and negative-coverage receipts come from [GBrain: projects-qwen36-bastet-lora; qwen36-bastet-lora-data, July 2026]. The truncation receipt comes from [GBrain: qwen36-bastet-lora-data, July 2026]. The other receipts come from [GBrain: projects/bastet-pipeline, July 2026] unless noted.
Positive examples alone teach suspicion. Fixed versions, clean contracts, and deliberate false-positive traps teach abstention. A target cut off mid-reasoning teaches the model to stop mid-reasoning. Fluent prose without a source span is not an audit finding.
Test the cheapest way to falsify the fine-tune
There is no universal dataset-size threshold in these receipts. Start with the option that could show that training is unnecessary.
| Approach | Use it for | Cost shape | Boundary |
|---|---|---|---|
| Prompt + RAG | First held-out test | $0 training; eval attention | Retrieval may close the gap |
| Synthetic-only SFT | Controlled bug variants | Training plus synthetic review | 9,369 rows were injected bugs |
| Agent-curated SFT | Grounded behavior | Generation hours plus review | Must pass blind-prompt and leakage gates |
| Full fine-tune vs LoRA | Green dataset | GPU-hours × rows × epochs | QLoRA execution is unverified |
| Do nothing first | Base model meets the bar | $0 | Stop before inventing a project |
Prompt + RAG is the first experiment. Give the base model repository context, require source spans, and score the held-out repositories. If it clears the bar, stop.
sft_grounded_v1 had 9,369 ChatML rows, all SolidiFI synthetic injected bugs ([GBrain: projects-qwen36-bastet-lora, July 2026]). That supports controlled coverage, not claims about production repositories, trap discrimination, or real-source localization.
The selector-plus-target-writer path produced blind ChatML prompts with sanitization, citation repair, and leakage validation. A separate re-export recorded 352 prompts: 283 train, 69 validation, and 80 skipped ([GBrain: projects-qwen36-bastet-lora, July 2026]). Skipped records document a refusal to manufacture a bad example.
QLoRA was planned but execution is unverified. A full run changes more weights and carries a larger compute and operations commitment; LoRA/QLoRA changes the adaptation shape and can be easier to iterate. Neither choice substitutes for clean data. We have no training-cost receipt and no evidence that a small clean run beats a large dirty one. That is a hypothesis, not a result.
Failures that blocked training
| Failure | Receipt | Impact | Response |
|---|---|---|---|
| Hash dedup defect | 287 records retained | Silent data loss | Audit count deltas |
| Artifact duplication | 8,005 rows, 302 artifacts, max 53 duplicates | Repetition looks like evidence | Dedup artifacts and prompts |
| Positive-only splits | 13 splits | Trains over-flagging | Add clean, fixed, trap cases |
| Truncation | 32% train rows | Breaks target behavior | Shorten or raise context |
| Grounding failure | Source localization missing | Findings cannot be checked | Halt below 70% recall |
The artifact receipt is not 8,005 independent observations. It can overweight a small source subset until the model learns repetition instead of the domain ([GBrain: projects-qwen36-bastet-lora, July 2026]).
Scale also broke the producer pipeline. 1,477 failures came from local-model degradation under GPU load: tool calls stopped and the model returned prose instead. The record identifies two fixable causes: GPU-load degradation and a SOUL.md instruction to “use file operations,” which made the model describe file work rather than call tools. The successful 21 cases used read_file; the fix required calls such as write_file and read_file. Afterward, all 63 regression tests passed ([GBrain: projects/bastet-pipeline, July 2026]).
The producer is part of the dataset. If it changes mode under load, output rows can mix actual work, prose about work, and malformed fallbacks. Log failure class before counting rows.
When not to fine-tune
Make the $0 call when any of these applies:
- Base model plus prompt/RAG passes the held-out repository evaluation.
- The prompt leaks a label, finding title, or answer.
- Source-localization candidate recall is below 70% ([GBrain:
projects-qwen36-bastet-lora, July 2026]). - You cannot count duplicates, negatives, truncation, and mappable labels.
- The generator fails under concurrency; reduce load, fix tool-use prompts, and rerun regressions first.
Eight workers were estimated at roughly 48 cases/hour. A 1,500-case run was estimated at 30–48 hours; selector work at 200–360 seconds/case and target writing at 80–160 seconds/case ([GBrain: projects/bastet-pipeline, July 2026]). These are pipeline-time, not training-time, receipts. They still show the trade: gates cost hours; contaminated production runs cost hours and produce misleading evidence.
Bottom line
The internal audit-data work found record loss, repeated artifacts, leaked labels, missing negatives, truncated targets, and broken source grounding before it produced a verified trained-model result. That is the result this report supports.
Run prompt + RAG first. Then pass dedup, labels, negative coverage, truncation, leakage, localization, and held-out-repository gates. Until leakage is zero, recall is at least 70%, negatives meet the target, and truncation is below 10%, keep the GPU-hours.
More on this decision, three ways to look at it:
Sources
- GBrain:
projects/bastet-pipeline, eight-stage pipeline, v0 dataset receipts, held-out repositories, dedup failure, tool-call failure/root cause, regression pass, throughput, and input constraints; facts recorded July 2026. - GBrain:
projects-qwen36-bastet-lora, synthetic and curated dataset lineage, negative ratio, duplicate-artifact receipt, grounding-pilot halt, candidate-recall gate, leakage constraints, and prompt export; facts recorded July 2026. - GBrain:
qwen36-bastet-lora-data, 137/39 train-validation split, over-4K rates, 32% train truncation, high-severity bias, and unmapped labels; facts recorded July 2026. - Recovery-plan path checked but not available:
/Users/kit/projects/qwen36-bastet-lora/docs/2026-06-05-production-sft-dataset-quality-recovery-plan.md(not found on 2026-08-21). The article relies on the GBrain receipts above.
Get the next verdict before it's everywhere.
One email when a new lab post or cost table ships. No spam, no confirmation step — unsubscribe anytime.