Building Specialized Models with ML Intern
Building Specialized Models with ML Intern Published October 8, 2026 A small version of the prompt rewriter included with Qwen Image 2.1 was needed. The official rewriter is a 9...
By Software Development Team
Building Specialized Models with ML Intern
Published October 8, 2026
A small version of the prompt rewriter included with Qwen-Image 2.1 was needed. The official rewriter is a 9B model that requires about 20 GB of memory and generates thousands of tokens before producing a paragraph. Only compressed versions of that model were available on the Hub.
After describing the desired system to ML Intern, a 0.8B version was produced the following day. The model runs on a CPU, returns valid output 99.7% of the time, and uses about one quarter as many tokens as the teacher model. The complete project, including having the 9B model label 8,797 example requests, cost USD 16 in compute.
Five additional models were built in the same way over the following days. Each project began as a message in HuggingChat with ML Intern enabled and ended as a public model on the Hub, with evaluation results documented in its model card.
ML Intern plans each project, requests a budget before spending anything, runs a small test before the main job, then trains, evaluates, and publishes the model on Hugging Face hardware.
Prompting ML Intern
The initial prompt receives most of the planning effort. The first prompt, for the citrus model, was about 450 words. By the sixth project, prompts were closer to 2,000 words because each project revealed requirements for the next one. All seven prompts are available in the yvrjsharma/ml-intern-prompts repository exactly as written.
A prompt generally begins with a one-line description of the idea and its purpose. It then specifies the dataset, base model, and training script. Information that has already been verified is placed under a section titled "Verified facts, do not re-derive". This lets the agent spend its budget on execution rather than rediscovering known details. For the camera-angle LoRA, this section identified the trainer that had recently added transparent-image support and the open GitHub issues that made an alternative trainer risky.
Two prompt requirements are particularly important:
- Request a baseline before training. The citrus prompt asks for the base model's zero-shot score on the same metric before training. Without this comparison, it is impossible to determine whether the trained model improved on its starting point.
- Add a smoke test with a verification step. For the image LoRAs, the test used 50 training steps, followed by a check that the saved weights had actually changed, before the full training run was authorized.
The prompt ends by listing the expected deliverables and setting a spending limit. It also defines what should be included in the model card and can include an instruction such as: "Cap total spend at USD 12 and ask me before exceeding it."
ML Intern starts every task with a zero-dollar budget and requires permission before executing paid jobs, so this limit is enforced. Without a specified budget, the agent proposes several approaches based on the project's size and asks which one to use.
The complete structure is not necessary for a first attempt. The original citrus brief did not include a verified-facts section, yet the resulting model more than tripled the accuracy of the Qwen3.5-2B model. The following examples show six projects built with ML Intern over several days.
1. A Model Specialized for a Field
A general vision model can describe a yellowing citrus leaf. Identifying whether the cause is a mite infestation or magnesium deficiency, then recommending biological and non-biological treatments, is considerably more difficult.
A training dataset was assembled from three sources hosted by the Project-AgML organization on the Hub. The resulting citrus-disease-vlm-instruct dataset contains 3,017 annotated images covering 21 pests, illnesses, nutritional deficiencies, and treatment approaches.
ML Intern fine-tuned Qwen3.5-2B on these examples and benchmarked the foundation model first. On 335 test photographs, the base model identified the correct problem 14.9% of the time. After two epochs on one A10G, the fine-tuned model reached 52.8%. Compute cost was about USD 1.90.
2. A Model That Draws a Specific Character
Image models know many characters, but Huggy, rendered in the flat style of Hugging Face brand assets, was not among them.
ML Intern was tasked with creating a LoRA for FLUX.2 klein base 4B, trained on 84 captioned drawings from the Chunte/huggy_for_training dataset.
The agent saved a checkpoint every 100 steps and generated the same prompts with each checkpoint. This made comparison straightforward. At step 200, Huggy first appeared fully consistent with the target style. From step 500 onward, the style began appearing in prompts unrelated to Huggy. The trained LoRA also works with the distilled Klein model at four steps. Compute cost was about USD 7.60.
3. Models That Add New Capabilities
Camera-angle LoRA
Camera-angle LoRAs are among the most popular community add-ons for earlier Qwen-Image models. They allow a user to provide an object image and request a view from a different angle, such as 45 degrees to the left. Several days after the release of Qwen-Image 2.1, no such add-on was available, so ML Intern was asked to build one.
ML Intern rendered 1,030 scanned household objects from Google Scanned Objects at 24 angles each. The result was 24,722 transparent images, generated in a CPU job that cost a few cents. The process then selected 461 objects for training and 40 for testing, producing 1,844 before-and-after training pairs distributed across 23 camera instructions.
Training ran for 2,000 steps in about 90 minutes on one A100, at a cost of approximately USD 3.75. The project took about half a day and involved 48 jobs, including jobs resubmitted after failures caused by missing packages or incorrect paths. Total compute cost was about USD 16.
Doodle-in LoRA
The Doodle-in LoRA accepts a photograph containing a magenta scribble and a short prompt naming an object. It replaces the scribble with that object while preserving the original lighting and composition.
Because no suitable dataset existed, the prompt specified how to create one. The process began with real photographs from Open Images. One object was removed using the LaMa inpainting model, and a scribble was drawn where the object had been. The untouched photograph served as the target.
ML Intern wrote and tested the pair-generation scripts in a CPU sandbox before running them as GPU jobs. The process also recorded the author and license of every source photograph. It produced 6,042 training pairs and a 160-pair test set. Forty test pairs came from 23 object classes excluded from training.
Before training, the base model was evaluated both alone and with the Viggle turbo LoRA. The process also verified that batched edits produced identical images, reducing evaluation costs. Training ran for 2,000 steps in 1 hour and 38 minutes on one A100, at a cost of approximately USD 4. Comparing checkpoints on 48 test pairs selected step 500.
When paired with the Viggle turbo LoRA at six steps, the LoRA detected the intended object in 67.5% of edits, with an average edit time of 4.7 seconds. Objects from the 23 unseen classes performed similarly to the other objects, at 65.0% versus 64.2%. The project took slightly more than a day and involved 59 jobs. Total compute cost was about USD 24.
4. Models That Fit on Smaller Devices
Pocket Rewriter
The Pocket Rewriter was built as a smaller alternative to the Qwen-Image 2.1 prompt rewriter. ML Intern first generated 8,797 short image requests using a small instruction model through Inference Providers. The prompts covered photographs, posters, logos, infographics, and other categories. About one third requested exact text in quotation marks, and many were written in languages other than English.
The 9B teacher rewrote all requests on one A100 in 2 hours and 37 minutes, at a cost of approximately USD 6.50. After quality filtering, 1,840 examples were selected for the training dataset.
Training the 0.8B and 2B students took 12 and 18 minutes respectively on an A10G, costing USD 0.75 for both. The 0.8B version is also available as an 812 MB GGUF file for CPU execution. The project took about 11 hours and involved 24 jobs. Total compute cost was approximately USD 16.
Agate-Preview-002-4step
Logolabs' Agate Preview 002 is a 260M-parameter text-to-image model that is small enough to run in a browser. However, its standard process requires 50 guided steps, equivalent to 100 network passes per image. The goal was to distill it to four passes.
The first training run cached 155,000 training images as latents, incorporated guidance into the model, and reduced the step count in stages from 16 to eight and then to four. The work ran on A100 GPUs. The four-step student outperformed the teacher running at four steps on GenEval and FID. It was then exported to ONNX for browser use. This run took about 13 hours and cost USD 22.
A second training run further improved the four-step student. ML Intern generated 24,000 additional image pairs with the teacher at 16 steps and fine-tuned the student against them for about an hour. GenEval increased from 0.509 to 0.536, compared with 0.563 for the teacher at 50 steps, while using only four steps. The browser version was exported again. The second run took about eight hours, bringing total compute cost for both runs to approximately USD 37.
Compute Costs
| Model | Base model | Compute |
|---|---|---|
| Citrus Doctor | Qwen3.5-2B | USD 1.90 |
| Huggy LoRA | FLUX.2 klein base 4B | USD 7.60 |
| Pocket Rewriter, 0.8B and 2B | Qwen3.5-0.8B and 2B | USD 16.05 |
| Viewpoint Orbit LoRA | Qwen-Image 2.1 | USD 16 |
| Doodle-in LoRA | Qwen-Image 2.1 | USD 24.30 |
| Agate 4-step, two runs | Agate Preview 002 | USD 37 |
| Total | about USD 103 |
These figures represent the GPU and CPU job charges reported for each session.
Building a Model
A project can begin with a model that should exist but is not yet available, together with a dataset that already exists or can be described precisely. A useful prompt identifies the base model, data, and training method, requests a baseline, includes a smoke test, and sets a small initial budget.
ML Intern can be accessed through HuggingChat by enabling ML Intern mode. The example prompts in the yvrjsharma/ml-intern-prompts repository provide starting points for similar projects.