what-i-learned-fine-tuning-a-small-model.md — 25 min read

2026-09-1625 min readManuel ParraInvestigation

What I learned fine-tuning a small model: mistakes you can avoid

What I learned directing coding agents through TabCrew fine-tuning: the task I set was too narrow, how I corrected it, and what the evaluation showed.

  • TabCrew
  • Fine-tuning
  • SFT
  • Evaluation
On this page · 8 sections
  1. 01What the model was being trained to do
  2. 02Training and evaluation contracts
  3. 03Training repair checkup before committing to a full run
  4. 04Three checkpoints, three different questions
  5. 05A useful title can change a preference
  6. 06Where the earlier RL experiments fit
  7. 07Conclusions
  8. 08Evidence and experiment notes

I have been working with coding agents to develop TabCrew, a new experimental extension for Google Chrome that uses AI to organize related tabs into named groups. The work goes like this: I plan the features, test the app, and turn the issues and errors I find into units of work. Subagents in Codex or Claude carry out that work, and then I review what they produced. Along the way, I have been learning more deeply how to train and evaluate models, especially small language models, or SLMs, fine-tuned for a specific task.

Throughout this article, I will be talking about two different components: the coding agents helping me investigate, experiment, and carry units of work through from start to finish, and the small model we are fine-tuning to organize tabs. Early on, I made a mistake in the direction of the training. The agents carried out the task as I had described it, without catching the gap between those instructions and what I intended. I own that mistake in the task definition. Correcting it meant preparing new examples and running additional training.

The issue was that I had prepared training data that taught the model to assign tabs to numbered groups, while the extension also needed relevant names for those groups. An answer could be valid JSON and still fail that requirement. I had defined the wrong output contract: the structure and information the extension expected to receive. That is what I want to walk through here: the work I delegated to the agents, examples of the artifacts we produced, and how I corrected my mistakes along the way. E02–E04

To start, TabCrew needs its groups to have names. The model should recognize which open pages belong together, assign every tab to a group, and give each group a useful name. The extension then needs a complete response it can apply through Chrome's tab-grouping APIs. Recognizing the topics is only part of that job; the answer also has to tell the extension which groups to create and where each tab belongs.

Consider an example session with pages about database backups and SQLite, places to go hiking over the weekend, and sourdough bread recipes. I might organize those tabs as “Database research,” “Weekend hike plans,” and “Sourdough recipes.” My first training target represented membership with numbers: these tabs belong to group 0, those to group 1, and the others to group 2. That can express a grouping, but it does not supply the names. The training run completed, yet I was teaching a narrower task than the extension needed. Completing the run did not establish that the first checkpoint could produce a usable response. E01, E02

The agents helped carry out the training pipeline, including annotation and synthetic-data preparation. I considered synthetic sessions a reasonable starting point for this experiment because the model works from tab titles and URLs to infer how the pages relate. That was a working assumption, though. Plausible synthetic sessions do not establish that a model handles real browsing well, and the data's limitations became part of what I needed to evaluate. E05, E07

Testing against the extension's complete contract made the mistake visible. The original checkpoint produced only one valid plan across 81 development sessions. I then directed a repair using 128 sessions with corrected answers to see whether we could recover the checkpoint. That repair produced 75 valid plans on the same comparison. I considered that a successful repair of much of the contract mismatch, with failures still left to investigate. E06, E09

After the pilot, I prepared a larger repair with 2,973 examples, close to the 3,030 sessions in the original training run. Both repairs started from separate copies of the original checkpoint; the larger run did not continue from the pilot. In the blind comparison, I chose the expanded repair in 51 of 81 cases. It produced 73 valid plans, slightly fewer than the pilot's 75. I consider that useful progress: I preferred its organization more often, while the pilot still led on validity. Those are different results, and the rest of the article explains what we did and how I evaluated them. E01, E08–E12

FIG. 01 A browser session, before organization
Browser session
Six tabs. Three activities.
01Postgres backup guide
02Restore a database snapshot
03Weekend ridge trail
04Trail weather forecast
05Sourdough starter care
06Whole-grain bread recipe

For the extension’s current review-and-apply flow, gateway boundaries, and beta limitations, see the TabCrew case study.

What the model was being trained to do

I started with Qwen3.5-0.8B, a small language model with about 800 million parameters. To adapt it to tab grouping, I used LoRA, which keeps the model's original weights fixed and trains a much smaller set of additional parameters. This let me work on the task without retraining the entire model. I used a LoRA rank of 8, a setting that controls the size of those trainable additions. E01

For the original training run, I used 3,030 browser-tab sessions and trained for one epoch. This means one pass through those training examples. The complete dataset contained 3,766 sessions in total. Of those, 367 were set aside for validation during development, and another 369 were reserved for the final test. Both sets were excluded from training. E01

One thing that amazes me is how much the saved example answer defines the task. Supervised fine-tuning, or SFT, trains against those answers. In this project, as I mentioned earlier, each training example connects a browser-tab session to a proposed grouping. The model reads the instructions and tab information, but the training loss (the error used to update the adapter) is calculated only on the expected answer. E01, E02

The prompt and padding tokens are excluded from that calculation. This is what it means to say they are masked in the loss. The model still reads the prompt as context; we are excluding it from what gets scored, not hiding it from the model. Padding tokens are filler added to make examples fit the length used in a training batch. They do not contain information about the tabs or the grouping, so we exclude them from the loss too. E01, E02

Why do that? Without those loss masks, the training objective would also include predicting the supplied prompt, including system instructions, and any padding tokens. For this task, I want the training to focus on producing the answer from the information I supply. Predicting filler would add a target that has nothing to do with organizing tabs. This choice focuses the training objective, but it does not guarantee that the answers I prepared teach the right task.

That is the part I understood later, after seeing the consequences of the training objective I had given my agents in the first place. If my saved answers omit group names, I am leaving those names out of what I train the model to produce. So even though the model was being trained to follow my examples, those examples were missing something the product contract needed. E02

Having the extension as a concrete use case is useful because it gives me an actual input and output contract to test in my local checkpoint comparisons. I can evaluate whether the answers fit what the product needs before deciding to integrate a checkpoint. The comparisons in this article are local development experiments; they do not demonstrate that the repaired model has been deployed in the extension.

What interests me about fine-tuning is that it feels like adapting a language model to your own use case, especially with these smaller models. Here, I am teaching one to organize tabs and return the answer my extension needs. In construction, I could prepare examples to teach a model to classify workers by trade, machines by equipment type, or companies by the services they provide. With suitable training examples, a small model may be able to handle that specific job. That is the idea I want to keep exploring: taking a model and training it for the work I actually need it to do.

These results give me a reason to keep investigating how small models can serve narrow tasks. The comparison here measures the validity and quality of the tab-grouping plans, not inference speed, latency, or cost. In separate local runs on modest hardware, I have observed generation throughput above 300 tokens per second. I consider that encouraging, but it is a preliminary observation outside the saved comparison in this article. I still need to record the hardware, model configuration, workload, and timing method and measure it more thoroughly. Tokens per second alone does not tell me how long a complete request takes or prove a cost advantage.

Training and evaluation contracts

After we finished the first training run, my first tests returned tabs assigned to numbered groups. Something was off: the answer did not match what I expected the extension to receive. I inspected the original student protocol, the instructions and answer format used to train the small model. It asked for assignments in this form:

Original student protocol · assignments onlyjson
{"assignments":{"1":0,"2":0,"3":1}}

If you read this answer, tabs 1 and 2 belong to group 0, and tab 3 belongs to group 1. It tells us which tabs belong together, but it does not define the groups' names or colors. The training input could also include synthetic intent, a description of the session's intended activity, when that field was present in an example. That extra input did not supply the missing names in the answer. E02

The gateway prompt used in this experiment expected a complete plan instead, with group definitions as well as assignments:

Gateway prompt · complete planjson
{
  "groups": [
    {"title":"Backup recovery","color":"blue"},
    {"title":"Weekend hike","color":"green"}
  ],
  "assignments":{"1":0,"2":0,"3":1}
}

Here is what the contract requires:

  • Each assignment points to a position in the group list, starting at 0. These positions are not Chrome's actual group IDs.
  • Every supplied tab must appear exactly once in the assignments.
  • Every referenced group must exist in the list.
  • Every defined group must contain at least one tab.
  • Group titles and colors must satisfy the contract. E03, E04

Chrome only allows a fixed set of tab-group colors: grey, blue, red, yellow, green, pink, purple, cyan, and orange. The model cannot invent any color it wants; its answer has to use one the API accepts. The group's title supplies its useful name.

The saved gateway input includes tab titles, URLs with query strings and fragments removed, and current-group context. The prompt treats that content as data to interpret, not instructions to follow. In the 81-case comparison, we did not supply synthetic intent, and the current-group fields were null. E03, E04

For future training, I want to prepare examples from actual browsing sessions and reviewed groupings, so the model can learn from how people organize their tabs as well as from synthetic examples. My expectation is that this could help improve the model over time. That is the next direction I want to explore, rather than a description of the data used in this comparison.

This is where I gave the agent the wrong task. I had treated the fine-tuning work as evidence that the model would be compatible with the application's more complete interface. But I had forgotten to check the training target against the response the application would actually consume. There was also an input mismatch: the training protocol sometimes supplied information that the gateway benchmark did not.

I do not think it makes sense to blame the agent for successfully completing the task I gave it. I had not been careful enough to verify that the task it was carrying out matched my intent. Agents can still make execution errors or misunderstand instructions. In this experiment, though, the mismatch was visible in the protocol from the start. Looking back, I suspect that checking it earlier could have saved me a couple of million tokens and a day of work.

Some of the initial criteria were still useful: every tab needed a group, and related tabs needed to be grouped together. The next unit of work was to create a small correction dataset with answers that followed the complete contract. I wanted to see whether additional training could repair the checkpoint so it produced the response I actually needed.

There was a second problem: my early evaluations did not make the gateway mismatch clear enough. A score that checked only which tabs were placed together could reward the model for putting two database pages in the same numbered group. But that answer could still omit the group's name and definition, leaving the extension without a complete plan to apply. The score was measuring one part of the task while missing a requirement that mattered to the product.

Our application also cleans up empty groups. That is useful when a user's changes leave a group empty, but a model creating an empty group in its proposed plan is a different issue: it has failed to follow the contract. I wanted that failure to receive a strong penalty, even if the application could clean it up afterward.

The corrected scorer therefore evaluates the unmodified response, before any application cleanup. Invalid plans receive a strict F1 score of zero and a reward of minus one. Putting related tabs together correctly cannot make up for an empty or undefined group, or a missing tab. The complete answer has to satisfy the contract. E04

FIG. 02 One task, two different interfaces

Input available to the model

Training could include synthetic intent

Titles, URLs, locale, current-group context, and synthetic intent when present.

Additional training context“Database work, a hike, and baking.”

The original target provided numeric assignments only. It explicitly excluded titles, names, and extra keys.

Original target shape

{
  "assignments": {
    "1": 0,
    "2": 0,
    "3": 1,
    "4": 1,
    "5": 2,
    "6": 2
  }
}
FIG. 03 A useful grouping still needs a valid plan

Backup recovery

  • 01Postgres backup guide
  • 02Restore a database snapshot

Weekend hike

  • 03Weekend ridge trail
  • 04Trail weather forecast

Sourdough

  • 05Sourdough starter care
  • 06Whole-grain bread recipe

Training repair checkup before committing to a full run

After figuring out what was happening, I assigned a new task to my agent: correct the examples so they matched the task the application actually asked for. Before committing to another full training run, I wanted to check whether a small repair could fix the problem.

We first audited a pool of 4,570 synthetic sessions. Those records already had numeric group labels, but none provided the complete named-group contract, so we needed to repair them. The pool included sessions used in the first training set; it was not 4,570 new examples. We used a model to review 30 samples. In 18, the grouping looked plausible. Nine raised questions about how broad or specific the groups should be, and three were set aside for further review and reannotation. E05

We started repairing a small subset of 128 sessions selected from the original training data. All were synthetic examples: 48 came from DeepSeek, 48 from IFM, and 32 from local Qwen. These were the sources of the original sessions. For the corrected answers, the new annotation requests used the same saved prompt as the gateway, with the input normalized to the gateway's format. We omitted synthetic intent and the old reference labels. E05

This time, the accepted answers contained both group definitions and assignments. The final audit recorded 128 unique accepted sessions with no contract failures. It also checked the saved records linking the corrected answers to their source sessions and annotation requests. So we were able to trace where each answer came from. The answers, though, had not yet been reviewed by a person. E06a

For the pilot, we continued training a separate copy of the original adapter for one epoch, with a batch size of four, a learning rate of 0.0002, and a seed of 42. The run completed 32 optimizer steps. I kept the original checkpoint so I could compare it with the repaired version. This was supervised training on the corrected answers. E06, E06a

These settings describe how the model learns from those examples:

Epochs

1

One epoch means one pass through the 128 training sessions. More epochs would expose the model to the same answers again. That can help it learn the task, but too much repetition can make it fit those particular examples without improving on new sessions.

Batch size

4

Batch size four means processing four sessions together for a training batch. In this run, 128 sessions divided into batches of four produced 32 updates. Batch size affects memory use and how many examples contribute to an update. Smaller batches usually produce more variable updates; larger batches combine more examples, but are not automatically better.

Learning rate

0.0002

The learning rate, 0.0002, controls the scale of the optimizer's changes to the adapter. A higher value can make training move faster, but it can also make updates too aggressive. A lower value makes smaller changes and may need more training. It is not a percentage improvement or an accuracy target.

Seed

42

The seed, 42, sets the starting point for the random choices controlled by the training code, such as shuffling examples. The number 42 has no special advantage. Keeping the seed fixed helps make comparisons more repeatable, although it does not guarantee identical results across hardware and software configurations.

Optimizer steps

32

The optimizer is the procedure that updates the trainable parameters using the gradients, which describe how the training loss changes with those parameters. An optimizer step is one update. Our 32 steps were 32 updates to the adapter, not 32 passes through the dataset.

Together, these settings control how often the model sees the examples, how they are combined, and how its parameters change. Changing them can change the result even with the same dataset. For this pilot, I wanted to see whether this small repair improved the answers before committing to a larger run; I had not established that these were the best possible settings.

Then we expanded the repair. We requested corrected answers for the pool of 4,570 synthetic sessions and obtained 4,484 annotations that passed the structural checks. The remaining 86 request outcomes were unresolved, so we left them out of this training run. Passing the structural checks meant an answer followed the required format and grouping rules. We still needed to check whether we could use the example for training. E07

That reduced the dataset further:

  1. 4,570

    requested sessions

  2. minus 86

    unresolved outcomes

    left out
  3. equals 4,484

    annotations that passed the structural checks

  4. minus 1,484

    1,484 annotations from older sources were left out pending an overlap audit.

    audit pending

    We needed to check those sources against the data reserved for evaluation across the different collections. Training on an evaluation example would make the later score less useful as a test of what the model had learned. These examples were set aside because that check was still pending, not because all 1,484 had been found to overlap.

  5. minus 27

    Another 27 annotations lacked the saved raw annotator output.

    left out

    We had to be able to check the processed answer against what the annotating model originally returned. Without that record, we could not independently verify the annotation in the same way, so we left those examples out too. E07

  6. equals 2,973

    corrected sessions for training

    admitted

The count was therefore 4,570 requested sessions, minus 86 unresolved outcomes, minus 1,484 awaiting the overlap audit, minus 27 missing raw outputs: 2,973 corrected sessions for training. I wanted the expanded run to use examples that passed these checks, rather than include every available answer just to keep the dataset larger. The exclusions reflected what we could verify for this run; they did not mean every excluded example had an incorrect grouping. E07

Of the 2,973 admitted sessions, 2,272 originally came from DeepSeek, 470 from IFM, and 231 from local Qwen. Those names describe the sources of the synthetic sessions. IFM supplied the corrected answers with the complete group definitions and assignments used for this repair. E07

The audit verified saved requests, receipt hashes, normalized outputs, and available raw outputs. Every admitted tab was assigned once to a defined, nonempty group with a nonblank name and an allowed color. It found no exact normalized-title overlap with 437 development sessions. Those checks reduce specific leakage risks; they do not rule out semantic paraphrases or shared synthetic patterns. E07

The expanded run also started from a separate copy of the original mixed-data adapter. It was not a continuation of the 128-session repair. It used one epoch, batch size four, seed 42, and learning rate 0.0002, completing 744 optimizer steps. Both repair checkpoints were retained for comparison. E08

I also had to separate a clean structural audit from evidence that the data resembled real browsing. Of the 2,973 admitted sessions, 2,908 (97.8%) contained placeholder domains. Titles often repeated the intended topic explicitly. The sessions contained 10 to 30 tabs and were classified as English or Spanish. An assistant's 12-example review found plausible boundaries and names, but that sample is neither human certification nor a population accuracy estimate. Real browsing includes vague titles, social profiles, overlapping activities, and much larger sessions. This dataset does not establish performance on those conditions. E07

Three checkpoints, three different questions

After finishing the repairs, I wanted to compare what we had produced. We now had three checkpoints: the original adapter, the repair trained on 128 sessions, and the expanded repair trained on 2,973. A checkpoint is the saved state of the model or adapter at a particular point in training. Keeping all three meant I could compare the answers instead of relying on how well a training run appeared to go.

We gave each checkpoint the same 81 development sessions and the same saved gateway messages. We also used the same XGrammar schema, which constrains the structure of the generated JSON. The decoding was greedy, meaning the model picked its most likely next token at each step, with a limit of 2,048 output tokens. Keeping those settings the same made the comparison easier to interpret. All three runs completed without execution errors. These were sessions we had already used during development, so I was using them to compare the repairs, not as the final test. E09, E10

I wanted to answer three questions: did the response follow the contract, did it group the tabs like the reference answer, and did I actually find the organization useful?

FIG. 04 81 development cases / four views of the result

Which plans obeyed the contract?

The pilot repair produced the most valid plans: 75/81 (92.6%), compared with 73/81 (90.1%) for the expanded repair. Validity checks the complete raw answer, not just the grouping.

Bar lengths share a zero-to-81 scale. This is a descriptive checkpoint comparison; it does not isolate a causal data-size effect.

See all numbers and scoring conditions
Frozen expanded-v1 evaluation · 81 development cases
MeasureOriginal128 repair2,973 repair
Raw-valid plans17573
Valid exact partitions05450
Strict pairwise F10.00970.87730.8207
Blind first choices12951
Shared top-choice credit13251
Valid shared top-choice credit03151

Shared credit requires exactly matching title strings and tab membership, ignoring group order, IDs, and colors. Three cases credit both repairs. Membership alone is insufficient because title usefulness was part of the review. Two first choices were contract-invalid. This was blind to the reviewer, not double blind.

Sources: matched evaluation and human utility review. Invalid full plans receive strict F1 = 0. Shared credit may exceed 81 when summed across models. E09–E12

The first question was the most direct. Could the extension use the complete plan? The original checkpoint produced only one valid answer. The small repair produced 75, and the expanded repair produced 73. That gave me a concrete result from correcting the training task: both repaired checkpoints were much more consistent about returning the contract I needed.

The second question needed a different check. We had a reference grouping for each session, and we compared the model's grouping with it. Pairwise F1 checks pairs of tabs: if two tabs belong together in the reference, did the model keep them together, and did it also put unrelated tabs together? The score combines how many of the expected pairings it recovered with how many of its proposed pairings were correct. In our strict version, an invalid complete plan gets zero before that grouping can earn credit.

An exact partition match is stricter about the grouping itself. It means every tab belongs with the same other tabs as in the reference. Changing the numeric group IDs does not matter. If the reference calls a group 0 and the model calls it 2, it is still the same grouping if its members match. The pilot matched 54 reference groupings exactly, and the expanded repair matched 50. Neither of these checks tells me whether the group names are useful.

So the small repair led on validity and agreement with the references. But when I reviewed the answers, I preferred the expanded repair more often. That was something I needed to understand before deciding which checkpoint I liked better. More training examples had not improved every result, and we had changed the answer format and input information as well as the amount of data.

A useful title can change a preference

For this part, I reviewed the answers as A, B, and C, with their order randomized. I could see the proposed organization, but not which checkpoint produced it or what score it received. That is what I mean by a blind comparison here. I was the reviewer, and the evaluation system still kept the mapping between each answer and its checkpoint. E10–E12

I wanted to choose the organization I would find useful, including the names. Across the 81 sessions, I picked the original checkpoint once, the pilot 29 times, and the expanded repair 51 times. Those choices were for evaluating the checkpoints; we did not use them as new training labels. E10–E12

Then we checked something that the first-choice count did not capture. In three cases, both repairs had produced the same named groups. Giving only one checkpoint credit depended on which equivalent answer I had selected. We therefore also calculated shared credit: both received credit when the group names and the tabs inside each group matched exactly. Group order, numeric IDs, and colors did not affect that check. This brought the pilot to 32 credits while the expanded repair stayed at 51. The original stayed at one. Because a case can give credit to two checkpoints, these counts add up to more than 81. E11

There were another 34 cases where an answer grouped the tabs the same way as my selected answer but used different names. We did not automatically count those as ties. The names were part of what I was reviewing. Two answers can put the same pages together while one gives me a clearer description of why they belong together. E11

My review also exposed a mistake in my own choices. Two answers I selected failed the complete contract: one from the pilot and one from the original checkpoint. I had liked what I saw, but that did not make the underlying response valid. When we counted shared credit only for valid plans, the totals became 0, 31, and 51. We kept the strict scores as they were. My preference could not excuse a missing part of the answer. E11, E12

This helped me separate what I was asking of the evaluation. I needed the automatic checks to catch broken plans, the reference comparison to measure agreement with the saved grouping, and my review to judge whether the organization and names were useful. Each was showing me something the other checks could miss.

Where the earlier RL experiments fit

Before these repairs, I had also worked on reinforcement learning, or RL. With supervised fine-tuning, we train against saved example answers. In RL, the model produces answers that receive scores, and those scores guide the training. The method we tried was GRPO, which uses groups of sampled answers and their relative rewards to guide updates.

The earlier RL runs on 18 sessions: valid answers, exact groupings, and what the result showed
RunValid answersExact groupingsResult
SFT only18 of 183not given in this note
SFT followed by GRPO18 of 184one additional exact match; mean reward change +0.056, resampled range -0.059 to +0.226improvement not established
GRPO alone15 of 18not given in this noteevery valid answer put all the tabs into one groupdegenerate
An earlier evaluation: a different experiment and baseline from the 81-session comparison above. Source: E13.

In an earlier evaluation with 18 sessions, the SFT-only model produced 18 valid answers and three exact reference groupings. The model trained with SFT followed by GRPO also produced 18 valid answers, with four exact groupings. That was one additional exact match, but I needed more than that small difference to say the RL stage had reliably helped. E13

We also looked at the change in mean reward, which was +0.056. To see how sensitive that result was to the evaluation sample, the analysis repeatedly resampled whole topic families and recalculated the difference. The resulting range was -0.059 to +0.226. It included both a decrease and an increase, and the evaluation had only three topic families. So I could report what happened in that experiment, but I could not confidently attribute a repeatable improvement to GRPO. E13

The run trained with GRPO alone showed a more obvious problem. Fifteen of its 18 answers were valid, but every valid answer put all the tabs into one group. It had satisfied the validity check without doing useful organization. That result made it clear why I needed to inspect the actual answers and how the reward was defined. E13

We later ran a small Muse test with three annotated sessions to check that the SFT and GRPO workflow could execute. Its validation change was too small to establish a reliable improvement. Both of these RL experiments came before the contract repairs discussed here. The repairs used supervised fine-tuning on corrected answers. E14

Conclusions

What I have learned is that the target answers in your training examples have to match the application contract and the problem you are trying to solve. I had directed my agents toward a training target that left out what the extension needed. Once we corrected the answers and aligned the input and output format, the number of valid plans went from one for the original checkpoint to 75 for the pilot and 73 for the expanded repair. I also preferred the organization and style of the expanded repair's answers in 51 of the 81 blind comparisons.

I consider that progress. It has given me more knowledge about fine-tuning, new perspectives, and clearer follow-up tasks. I want to understand the remaining invalid answers and check whether tighter generation rules or targeted corrections can fix them. I also want to test the model on more realistic browsing sessions and measure its performance. I plan to share those results in a later update, once the model is in production and I have completed the experiments.

Working through this with agents has made me pay more attention to what I actually delegate. I can ask for annotation, training, and evaluation and get a complete pipeline going, but I still need to check that the training matches the outcome we are looking for and what the product needs.

That is what I will carry into the next task: define a narrow, specific problem around what my application needs, check the examples before committing to a training run, and look at what the model actually produces. I am still learning how to do all of that well, but now I have hands-on experience with what to check before asking an agent to start another training run.

Evidence and experiment notes

These notes summarize the saved experiment records behind the claims above. The underlying raw records are not published with this article. The notes describe the evidence and its limits; they are not links to publicly downloadable source files.

E01

Mixed-data training

3,030 training sessions; 367 validation; 369 reserved final; rank-8 LoRA; completion-token loss; one epoch. This is the original mixed adapter, not the earlier generator-label baseline.

Back to the article

E02

Original student protocol

Assignments only; names explicitly excluded; synthetic intent inserted when present. The first 29 lines establish the prompt and input contract.

Back to the article

E03

Saved gateway prompt

Named definitions plus assignments, normalized URLs, untrusted inputs, complete coverage, nonempty groups, valid names and colors. This is the saved experimental prompt; no claim that it is the latest deployed prompt.

Back to the article

E04

Contract diagnosis and strict scoring

Raw response validity is a hard gate. Assignment-only learning is not complete-contract compatibility. Current-group context was null for these development cases. Historical cleaned scores are superseded by the raw-contract correction at the top.

Back to the article

E05

Pilot data audit

4,570 candidate sessions; all numeric partitions; synthetic intent mismatch; 30 assistant-reviewed samples; pilot selection of 48/48/32. This document contains historical progress states; completed receipts E06 supersede its pending-training statements.

Back to the article

E06

Completed pilot training

completed=true; 128 sessions, 32 optimizer steps, one epoch, batch four, seed 42, learning rate 0.0002; original adapter continuation, no RL.

Back to the article

E06a

Pilot label audit

128 accepted unique sessions; zero contract failures; provenance verified; human_reviewed=false.

Back to the article

E07

Expanded admission and limitations

4,570 requests; 4,484 structurally validated; exclusions 1,484 + 27; 2,973 admitted; 2,908 placeholder-domain sessions; source origin is distinct from corrected-label provider; overlap and realism limitations.

Back to the article

E08

Completed expanded training

completed=true; 2,973 sessions, 744 optimizer steps, one epoch. It starts from the original mixed adapter, not the pilot adapter.

Back to the article

E09

Matched development scores

81 cases each; validity 1/75/73; exact partitions 0/54/50; strict F1 0.009676343/0.877277589/0.820655460; zero execution errors.

Back to the article

E10

Frozen comparison configuration

Identical saved gateway messages; same XGrammar schema, greedy argmax, 2,048 output limit; hashes for three adapters and cases; fresh runs; blind not double blind; sealed_test_used=false.

Back to the article

E11

Shared human-utility credit

First choices 1/29/51; shared credit 1/32/51; valid shared credit 0/31/51. Exact title strings plus membership, ignoring group order/IDs/colors. One reviewer over development cases.

Back to the article

E12

Human review analysis

81 reviewed; winner counts verified; two chosen-invalid cases; mapping_matches_rendered_outputs=true; training_labels=false. No individual review text or private session data are copied into this package.

Back to the article

E13

Historical generator-label SFT/GRPO

18-session earlier final evaluation; SFT exact 3, SFT+GRPO exact 4; reward delta +0.056 with descriptive family-block interval [-0.059,+0.226]; GRPO-only valid outputs all degenerate. Different experiment and baseline from E09.

Back to the article

E14

Historical three-session Muse smoke test

Actual task annotations and actual SFT/GRPO updates; nine already-exposed validation cases; reward change +0.004; no reliable RL or generalization claim.

Back to the article