Posts

  • Order still matters

    The first post in this series found that field order sometimes changes output quality, and it ended with a promise to check whether that holds on a small model that doesn't think, from a family other
  • Declared, not delivered

    Field order sometimes changes output quality. That was the finding in the first post in this series, and the follow-up I said I would run was whether the effect holds on a small model that doesn't thi
  • A number that looks fine

    If you have measured the quality of LLM output at any scale, you have probably built one of these. You write a rubric of a few criteria, you have a strong model score each one from 1 to 5, and you com
  • Thinking your way out

    In one of the levels the model generated, the final boss is weak to an item the player never has a chance to pick up. Every field reads as plausible and nothing is malformed, but the weakness points a
  • Working as intended

    QA is being cut on the theory that AI can test now. The theory is mostly right. But the role bundles two jobs under one title, and the cut can't tell them apart. One of them AI just finished. The othe
  • Preview is preview

    A preview model is a deal. You get something good early, you accept the rough edges like throttling, inconsistency, and surprise behavior, and you accept the clock. The deal is worth taking when the p
  • Nothing to hide behind

    A prompt firewall is cheap to run in front of a language model because the language model is slow. A sanitization call that adds 80 ms disappears into the second or two a model takes to respond. You c
  • Autoregressive schemas

    LLMs are good at producing structured output from messy input. That property is core to agentic systems and to tasks like structured data extraction. Both are work I do day to day, at my job and on si

subscribe via RSS