Decision models return a single choice or probability, like “yes,” “spam,” or a 1-10 score, in one step instead of generating a full written response, making them far faster and cheaper for simple classification tasks.
OpenAI launched its Decisions API last week, running on GPT-6 Luna, charging $0.10 per million input tokens with no separate charge for the output since none is generated as text. Separately, open-source toolkit Unsloth released a way to fine-tune any existing small model, like Qwen3.8 or Gemma 4, into this same decision-making format, and says it raised one small model’s accuracy from 20.7% to 74.3% on three decision benchmarks using just 4GB of GPU memory.
These tools suit high-volume, repetitive choices, like sorting support tickets or flagging spam, where a full written explanation would be wasted and the person doesn’t need to understand the model’s reasoning.
Key Capabilities:
- No output tokens: Returns a direct answer instead of generating explanatory text, cutting cost and latency.
- DIY option: Unsloth’s approach lets developers convert their own small model into a decision model, not just use OpenAI’s.
- Low hardware bar: Unsloth’s fine-tuning runs on just 4GB of GPU memory.