Skip to content
TP
← Back to the lab
Ongoing2026

Shell Command Model

A small model that turns plain requests into correct shell commands across macOS, Linux and Windows.

PythonSmall language modelsEval harnessONNX

The question

How small can a model get before it stops writing correct tar flags?

How it works

PROMPTplain request01ROUTEdetect target OS02GENERATEcandidate command03VERIFYdry-run and lint04SCOREexact-match eval05

Where the time went

effort
  • Evaluation harness50%
  • Model comparison30%
  • Safety and dry-run20%

How it progressed

relative momentum over the build

What I found

Smaller than expected, if the eval is strict. Most tiny tool-calling models are graded far too generously.

Next experiment

PocketLink

$./next-step

Most of these started as a question I could not answer by reading.

Send me a hard one.