The Search For The Right AI Model Led To Making One
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Search For The Right AI Model Led To Making One on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A Hugging Face contributor reports using the ML-intern agent to build and publish seven custom models over several days. The account gives results and compute costs for selected projects, but the figures are self-reported and have not been independently verified.

A Hugging Face contributor says they used the platform’s ML-intern agent to build and publish seven custom models over several days, including a compact prompt rewriter and a citrus disease classifier. The account describes an agent-assisted workflow for planning and running model training, but its reported results and costs are the contributor’s figures, not an independent evaluation.

The first project addressed a 0.8-billion-parameter prompt rewriter intended as a smaller alternative to the rewriter included with Qwen-Image 2.1. The contributor said that model has 9 billion parameters, requires about 20 GB of memory and can generate thousands of tokens before producing a paragraph. They reported that the smaller model produced valid output 99.7% of the time and used about one-quarter as many tokens as its teacher. The reported compute cost, including labeling 8,797 example requests with the larger model, was about $16.

Another project fine-tuned Qwen3.5-2B to identify citrus pests, diseases and nutritional deficiencies in images. According to the contributor, its dataset included 3,017 annotated images across 21 categories. On 335 test photos, they reported that the base model identified the correct problem 14.9% of the time, compared with 52.8% for the fine-tuned model after two training epochs on one A10G GPU. The contributor put the compute cost at about $1.90.

The account also describes a character-generation LoRA trained on 84 captioned drawings and a camera-angle LoRA for Qwen-Image 2.1. For the latter, the agent generated 24,722 transparent images of scanned household objects across 24 angles. Training took about 90 minutes on one A100; the contributor said the project cost about $16 in compute, including failed jobs that had to be resubmitted. They said model cards and evaluations were published on Hugging Face. The source gives detail on selected projects, not a full account of all seven.

At a glance
reportWhen: Reported over several days; the source…
The developmentA Hugging Face contributor says the ML-intern agent helped build and publish seven custom models, with reported compute costs ranging from about $1.90 to $16 for described projects.
At a glance
reportWhen: Reported last week; the projects were b…
The developmentA Hugging Face contributor says an AI agent called ML-intern helped plan, train, evaluate and publish seven custom models on the Hub over several days.

A Shorter Route to Custom Models

The account offers a concrete example of how an AI agent might lower the hands-on coordination needed to customize a model. The contributor said they started each project in HuggingChat with ML-intern enabled. The agent proposed a plan, asked for spending approval before paid work, ran a small test and then handled training, evaluation and publication using Hugging Face hardware. If a prompt did not specify a budget, the agent presented options and asked the user to choose.

The reported figures suggest that some experiments can be run for a modest amount of compute spending. They do not establish the full cost of making a model: the source does not account for all time spent preparing data, writing instructions, checking outputs or deciding whether the results are useful. Nor does it show whether the same results would be typical for other users, tasks or budgets.

The projects also illustrate why a model’s score needs a comparison baseline. The citrus result is presented against the untuned model on the same test set, rather than as a standalone accuracy claim. The account also reports a limitation: later character-model checkpoints began affecting prompts unrelated to the character, which the contributor described as a sign of overfitting or overly broad style effects. A successful training run is not, by itself, evidence that a model will behave reliably beyond its intended use.

Amazon

AI model training GPU

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From a Missing Model to Seven Projects

The initial request came from the contributor’s search for a smaller prompt rewriter for Qwen-Image 2.1. They said they found compressed versions of the 9-billion-parameter model on Hugging Face but no smaller alternative for their needs, so they used the larger model as a teacher to help create a compact one.

According to the contributor, their instructions became more detailed as the projects progressed, growing from about 450 words for the first project to nearly 2,000 for the sixth. The prompts specified the dataset, base model and training script, and asked for a baseline, a small test run and a spending cap. The contributor said all seven prompts are available in a public GitHub repository.

That process is a description of one person’s experience, not a controlled comparison of the agent against manual model development. The published model cards, evaluations and prompts may let others inspect the work, but they do not by themselves establish that the reported gains or costs will carry over to other projects.

“Also report the base model’s zero-shot score on the same metric before training so we can see the gain.”

— The Hugging Face contributor

Amazon

machine learning model fine-tuning kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Reliable Are the Reported Results?

The figures are self-reported, and the supplied account does not provide independent replication or full evaluation protocols for every project. It does not explain how the 99.7% valid-output rate was measured, whether test images were independently reviewed, or how data quality was checked across the projects. The detailed descriptions cover selected work rather than all seven models.

The account also leaves open how well the models would perform on new or different data, or under different hardware and spending limits. The compute estimates do not include a full accounting of data preparation, prompt writing or review time. The contributor described the camera-angle training run and character-model behavior, but the source does not provide enough comparable results to assess the agent’s performance across the full set of projects.

Amazon

AI image classification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Published Models Invite Further Checks

The contributor says the seven models and their evaluations are available through Hugging Face, while the project prompts are available on GitHub. Those materials give interested users a way to examine the examples and compare them with their own requirements; the source does not announce a specific next release or independent review.

Further checks across users, datasets and tasks would help show whether the reported costs and performance gains are typical. For later projects, the contributor’s described process calls for setting a spending limit, running a small test and recording the base model’s score before training. Whether those checks are sufficient will depend on the application and on the quality of its data and evaluation.

Amazon

prompt rewriting AI tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the ML-intern agent reported to have done?

The contributor says it helped plan jobs, request spending approval, run test and training work, evaluate models and publish them using Hugging Face hardware. This describes one contributor’s account, not an independent assessment of the agent.

What were the reported results for the citrus classifier?

The contributor reported that the fine-tuned model identified the correct citrus problem in 52.8% of 335 test photos, compared with 14.9% for the base model. They said it was trained for two epochs on one A10G GPU using 3,017 annotated images across 21 categories.

How much did the described projects cost?

The contributor reported about $1.90 in compute for the citrus classifier and about $16 for both the prompt rewriter and camera-angle LoRA projects. These are reported compute costs, not a full accounting of time or every possible expense.

Have the results been independently verified?

The supplied account does not report independent replication. It also leaves some evaluation details unspecified, including how the prompt rewriter’s valid-output rate was measured and how the test images were reviewed.

Where can readers inspect the work?

The contributor says the models and evaluations are on Hugging Face and that the project prompts are in a public GitHub repository. The source does not give enough information to independently confirm the completeness of every project description.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Cloud Services Surges In Global Coverage

Cloud services are experiencing a sharp increase in global media mentions, with GDELT reporting 26 times the usual coverage in recent days.

AirPods Pro 3 Or AirPods 5? I Compared Them For Weeks, And It’s Surprisingly Close

Search and coverage interest is rising around an AirPods Pro 3 and AirPods 5 comparison, but the reason for the spike is unconfirmed.

How Artificial Intelligence Is Enabling Multilingual Understanding

Google announces support for over 300 languages, real-time speech translation in 70 languages, and new models to improve multilingual understanding worldwide.

Stepping Inside The Design: How VR Is Advancing USACE Projects [Image 3 Of 7]

Virtual reality is now being used to advance USACE projects, allowing for immersive design reviews and improved project planning.