ej - An 11MB local model for typed decisions in one pass

by•
I built ej for a narrower problem than general classification: given some state and a set of decision questions whose options can be defined at runtime, return a probability distribution for each question in one pass. It handles choice, yes/no, and ordered scores, runs on CPU, needs no token generation, and ships as one 11.4 MB model file. The repo includes training, evaluation, adaptation, packed weights, and the benchmark harness.

Add a comment

Replies

Best
I built ej because I wanted to see how far a very small decision model could go without hiding the tradeoffs. The road there wasn’t straight: 🔀 It started in Rust. The first version was a Rust engine. Somewhere along the way the research loop outgrew it, and ej became Python. Being able to iterate quickly turned out to matter more than the language I’d started in. 📦 “11 MB” was a lie, at first. Early on, 11 MB was only a bit-level count of the parameters. The file you actually downloaded was much bigger. So I wrote my own packed format to make the number real. Today it’s one 11,384,312-byte file with no pickles and a SHA-256 check on every section. ⚖️ A licence check sent me back to the start. Late in the project, one training dataset turned out to be non-commercial. I removed it, retrained the encoder and refit the whole model. Some accuracy was lost, and it was worth it. 🎯 I set a target and missed it. I aimed for 60% accuracy on workflows ej had never seen. The honest result is about 42%, and it’s the first thing I point people to. 🧪 My best idea failed my own test. A small model that reads the text and the options together looked great in full precision. Quantised to fit, its gain fell below the bar I had set in advance, so I dropped it. Several other “improvements” went the same way once I trained each one several times and found the gains were within run-to-run noise. A lot of the real work was the unglamorous part: packing the model, deterministic inference, offline loading, measuring calibration on each test suite instead of promising it, and a benchmark where contamination and uncertainty are visible. The result is useful in some settings and clearly weak in others, especially on unseen workflows. I’ve tried to make those limits easy to inspect instead of smoothing them over. I’d love feedback on two things in particular: the benchmark design, and whether the current size/accuracy tradeoff feels useful in real systems.