Applied LLMs: What We Learned From a Year of Building
A long, specific write up from six practitioners on what actually worked and failed when shipping language model products, covering evaluation, retrieval, operations and team structure. Rare in that it reports the failures.
Guide
AI & Machine Learning
Yan, Bartolo, Fu and others
3 to 4 hours
Advanced
Free
What it is
A collaborative report from engineers who spent a year building with large language models, organised into tactical, operational and strategic sections.
Why it is unusually valuable
It is specific about failure. It says which evaluation approaches did not work, why retrieval quality mattered more than model choice, why fine tuning was usually the wrong first move, and how much of the effort went into data and evaluation rather than modelling. Almost nothing else in the field is written with that candour.
The sections to read first
The evaluation material, because building a reliable evaluation set is the single highest leverage thing most teams are not doing, and the section on when to fine tune, because it will probably save you a month.
Best for: teams beyond the prototype stage who are discovering that shipping is the hard part.
Ready to start?
Opens on Yan, Bartolo, Fu and others in a new tab.
Stuck on something specific?
Writing only gets you so far. If you want an answer to your situation rather than the general case, book a session and we will work through it together. Every session is free; a few slots open each week.
Follow along
New writing, resources and project ideas land here first.