This software tests whether LoRA fine-tuning helps one small model (Qwen2.5-Coder-1.5B) explain Rust compiler errors, using a 12-record dataset.
Note
Status: Research prototype; dataset checks pass and the negative result is preserved.
- Checks dataset fields and keeps Rust error families apart across data splits (dataset, validator).
- Compares model answers with one deterministic rule baseline (baseline); the recorded run showed no gain (report, evidence). On the three held-out records the rule baseline scores 1/3 exact strategy match, ahead of both model variants (0/3 each); its rule table is hand-written and includes an entry for a held-out error code, so it is a plumbing floor, not a tuned competitor (details in the report).
- Can tune a small set of model weights with LoRA (Low-Rank Adaptation) using Apple's MLX toolkit on Apple Silicon (training, evaluation).
- Checks Rust examples with the Rust compiler and records evaluation evidence (compile checks).
You need Git and Python 3 (CI uses Python 3.11; download Python). The data checks and baseline use the standard library and need no model weights, GPU, API key, or MLX installation. Start with:
git clone https://github.com/rustfuture/model-adaptation-lab.git
cd model-adaptation-lab
python3 scripts/validate_dataset.py
python3 scripts/validate_dataset_v2.py
python3 scripts/baseline.pyThe validators print dataset split summaries; the baseline prints its score on the test records. These checks do not train or download a model. On Windows, use your Python 3 command (for example py -3) in place of python3.
For the Rust snippet checks, also install a Rust toolchain with Cargo (rustup). The shell safety check below requires Bash; these commands use a macOS/Linux shell.
python3 scripts/validate_rustc_snippets.py
python3 scripts/validate_rustc_snippets_v2.py
python3 scripts/verify_evidence.py --write-manifest /tmp/evidence-metadata.json
diff -u evidence/metadata.json /tmp/evidence-metadata.json
python3 -m unittest discover -s tests -p 'test_*.py'
./tests/test_script_safety.shModel training needs an Apple Silicon computer, MLX, mlx-lm, and the model weights, which are not included here:
python3 scripts/prepare_mlx_data.py
scripts/prepare_mlx_model.sh
run_manifest="$(set -o pipefail; scripts/train_mlx_lora.sh | tee /dev/stderr | sed -n 's/^Training complete\. Run manifest generated at: //p')" &&
python3 scripts/evaluate_mlx.py --manifest "$run_manifest"The v2 runner executes its full pipeline and reports when model weights are unavailable: ./scripts/run_v2_experiment.sh.
- The validators check required dataset fields and keep each Rust error family in one split. They also compile source examples and check for the expected compiler error.
- Preparation scripts turn examples into prompts and answers for model training and evaluation.
- On Apple Silicon, MLX can tune a small set of model weights with LoRA, a method called Low-Rank Adaptation; run data stays in isolated temporary folders.
- The evaluators compare saved model answers with expected answers. The v2 evaluator also formats generated Rust code and checks whether it compiles.
- Evidence scripts recalculate file digests and metrics from saved outputs.
See docs/reference.md for experiment details, results, file safety rules, and the repository map.
The v2 evaluator checks whether generated Rust code compiles. Passing that check does not show that the code behaves correctly or that the explanation is right.
| Level | Meaning here | Established by |
|---|---|---|
| Compiles | The snippet builds as a standalone Rust 2021 binary | validate_rustc_snippets_v2.py, v2 evaluation |
| Behaviorally correct | The fixed code passes behavior tests | Not claimed; the dataset has no behavior tests |
| Semantically correct | The diagnosis and fix are right | Not claimed; keyword and exact-match proxies only |
Exact text match and keyword coverage are narrow measures. They do not establish semantic correctness.
The CI workflow in .github/workflows/ci.yml runs these commands:
python -m pip install --disable-pip-version-check --no-input "pyflakes==3.2.0"
python -m pip check
pyflakes scripts/
python3 scripts/validate_dataset.py
python3 scripts/validate_rustc_snippets.py
python3 scripts/validate_dataset_v2.py
python3 scripts/validate_rustc_snippets_v2.py
python3 scripts/validate_dataset.py --write-manifest "$RUNNER_TEMP/dataset-validation.json"
diff -u evidence/dataset-validation.json "$RUNNER_TEMP/dataset-validation.json"
python3 -m unittest discover -s tests -p 'test_dataset_contract.py'
python3 -m unittest discover -s tests -p 'test_v2_provenance.py'
python3 scripts/baseline.py
python3 scripts/prepare_mlx_data.py
git diff --exit-code -- data/mlx
python3 scripts/verify_evidence.py --write-manifest "$RUNNER_TEMP/evidence-metadata.json"
diff -u evidence/metadata.json "$RUNNER_TEMP/evidence-metadata.json"
./tests/test_script_safety.shMore test commands
The tests check script lint, dataset fields and splits, Rust compilation, baseline scoring, prepared data, saved evidence, and filesystem path safety.
- The recorded run showed no quality gain. This single small experiment says nothing general about LoRA or Qwen models.
- The recorded run trained on dataset SHA-256
bd488f58...; the committeddata/rust_errors.jsonlis now505ba845...because five records'codefields were edited on 2026-09-15 (commitcf0e452).data/mlx/*regenerated from the current file is therefore not byte-identical to the recorded run's training data. See evidence/README.md. - The historical adapter is not included, so its full training run cannot be repeated from this checkout.
- The small test set and hand-written keyword measure cannot establish significance or semantic correctness.
- The loss pattern fits overfitting, but the experiment does not establish the cause.
- MLX training only runs on Apple Silicon. Colab checks the data, baseline, and evidence; it does not run MLX training.
MIT; see LICENSE.
