Skip to content

About

This software tests whether LoRA fine-tuning helps one small model (Qwen2.5-Coder-1.5B) explain Rust compiler errors, using a 12-record dataset.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Model Adaptation Lab

model-adaptation-lab project overview

This software tests whether LoRA fine-tuning helps one small model (Qwen2.5-Coder-1.5B) explain Rust compiler errors, using a 12-record dataset.

CI License: MIT Open validation in Colab

Note

Status: Research prototype; dataset checks pass and the negative result is preserved.

  • Checks dataset fields and keeps Rust error families apart across data splits (dataset, validator).
  • Compares model answers with one deterministic rule baseline (baseline); the recorded run showed no gain (report, evidence). On the three held-out records the rule baseline scores 1/3 exact strategy match, ahead of both model variants (0/3 each); its rule table is hand-written and includes an entry for a held-out error code, so it is a plumbing floor, not a tuned competitor (details in the report).
  • Can tune a small set of model weights with LoRA (Low-Rank Adaptation) using Apple's MLX toolkit on Apple Silicon (training, evaluation).
  • Checks Rust examples with the Rust compiler and records evaluation evidence (compile checks).

Quick start

You need Git and Python 3 (CI uses Python 3.11; download Python). The data checks and baseline use the standard library and need no model weights, GPU, API key, or MLX installation. Start with:

git clone https://github.com/rustfuture/model-adaptation-lab.git
cd model-adaptation-lab

python3 scripts/validate_dataset.py
python3 scripts/validate_dataset_v2.py
python3 scripts/baseline.py

The validators print dataset split summaries; the baseline prints its score on the test records. These checks do not train or download a model. On Windows, use your Python 3 command (for example py -3) in place of python3.

Optional compiler and evidence checks

For the Rust snippet checks, also install a Rust toolchain with Cargo (rustup). The shell safety check below requires Bash; these commands use a macOS/Linux shell.

python3 scripts/validate_rustc_snippets.py
python3 scripts/validate_rustc_snippets_v2.py
python3 scripts/verify_evidence.py --write-manifest /tmp/evidence-metadata.json
diff -u evidence/metadata.json /tmp/evidence-metadata.json
python3 -m unittest discover -s tests -p 'test_*.py'
./tests/test_script_safety.sh

Model training needs an Apple Silicon computer, MLX, mlx-lm, and the model weights, which are not included here:

python3 scripts/prepare_mlx_data.py
scripts/prepare_mlx_model.sh
run_manifest="$(set -o pipefail; scripts/train_mlx_lora.sh | tee /dev/stderr | sed -n 's/^Training complete\. Run manifest generated at: //p')" &&
python3 scripts/evaluate_mlx.py --manifest "$run_manifest"

The v2 runner executes its full pipeline and reports when model weights are unavailable: ./scripts/run_v2_experiment.sh.

How it works

  • The validators check required dataset fields and keep each Rust error family in one split. They also compile source examples and check for the expected compiler error.
  • Preparation scripts turn examples into prompts and answers for model training and evaluation.
  • On Apple Silicon, MLX can tune a small set of model weights with LoRA, a method called Low-Rank Adaptation; run data stays in isolated temporary folders.
  • The evaluators compare saved model answers with expected answers. The v2 evaluator also formats generated Rust code and checks whether it compiles.
  • Evidence scripts recalculate file digests and metrics from saved outputs.

See docs/reference.md for experiment details, results, file safety rules, and the repository map.

Correctness Levels

The v2 evaluator checks whether generated Rust code compiles. Passing that check does not show that the code behaves correctly or that the explanation is right.

Level Meaning here Established by
Compiles The snippet builds as a standalone Rust 2021 binary validate_rustc_snippets_v2.py, v2 evaluation
Behaviorally correct The fixed code passes behavior tests Not claimed; the dataset has no behavior tests
Semantically correct The diagnosis and fix are right Not claimed; keyword and exact-match proxies only

Exact text match and keyword coverage are narrow measures. They do not establish semantic correctness.

Tests

The CI workflow in .github/workflows/ci.yml runs these commands:

python -m pip install --disable-pip-version-check --no-input "pyflakes==3.2.0"
python -m pip check
pyflakes scripts/
python3 scripts/validate_dataset.py
python3 scripts/validate_rustc_snippets.py
python3 scripts/validate_dataset_v2.py
python3 scripts/validate_rustc_snippets_v2.py
python3 scripts/validate_dataset.py --write-manifest "$RUNNER_TEMP/dataset-validation.json"
diff -u evidence/dataset-validation.json "$RUNNER_TEMP/dataset-validation.json"
python3 -m unittest discover -s tests -p 'test_dataset_contract.py'
python3 -m unittest discover -s tests -p 'test_v2_provenance.py'
python3 scripts/baseline.py
python3 scripts/prepare_mlx_data.py
git diff --exit-code -- data/mlx
python3 scripts/verify_evidence.py --write-manifest "$RUNNER_TEMP/evidence-metadata.json"
diff -u evidence/metadata.json "$RUNNER_TEMP/evidence-metadata.json"
./tests/test_script_safety.sh
More test commands

The tests check script lint, dataset fields and splits, Rust compilation, baseline scoring, prepared data, saved evidence, and filesystem path safety.

Limitations

  • The recorded run showed no quality gain. This single small experiment says nothing general about LoRA or Qwen models.
  • The recorded run trained on dataset SHA-256 bd488f58...; the committed data/rust_errors.jsonl is now 505ba845... because five records' code fields were edited on 2026-09-15 (commit cf0e452). data/mlx/* regenerated from the current file is therefore not byte-identical to the recorded run's training data. See evidence/README.md.
  • The historical adapter is not included, so its full training run cannot be repeated from this checkout.
  • The small test set and hand-written keyword measure cannot establish significance or semantic correctness.
  • The loss pattern fits overfitting, but the experiment does not establish the cause.
  • MLX training only runs on Apple Silicon. Colab checks the data, baseline, and evidence; it does not run MLX training.

License

MIT; see LICENSE.

About

This software tests whether LoRA fine-tuning helps one small model (Qwen2.5-Coder-1.5B) explain Rust compiler errors, using a 12-record dataset.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages