New API for Probability-Based Classification
OpenAI has officially launched its new v1/decisions API in public beta, providing developers with a specialized tool for classification and sorting tasks. Unlike standard chat-based endpoints, this API returns typed answers, probabilities, and confidence scores directly from text and image inputs. According to Generative AI News, the endpoint is designed to significantly outperform the standard Responses API in speed, with OpenAI claiming it is approximately 10 times faster for specific classification workloads.
The API supports three distinct query types: predicate, which returns the probability that a specific condition is met (e.g., detecting a scratch in an image); choice, which selects from a set of predefined options; and score, which assigns values to ordered stages. By receiving these probabilities directly, developers can establish their own thresholds for decision-making, moving away from the need to interpret free-form text responses from a general-purpose model.
Pricing Structure and Operational Constraints
A notable feature of the Decisions API is its pricing model, which focuses exclusively on input costs. OpenAI has set the rate at $0.10 per 1 million tokens. Critically, the official documentation specifies that there are no charges for output tokens or cache reads and writes. While regional processing surcharges and multipliers for long inputs still apply, the elimination of output billing provides a predictable cost structure for high-volume classification tasks.
Currently, the API is limited to the gpt-6-luna model. Developers should also note that image inputs must be provided as inline base64 strings, as the API does not currently support file IDs or direct URL references. OpenAI expects the service to move to general availability within a few weeks. For developers requiring free-form text or custom JSON structures that fall outside the scope of classification, the company notes that its existing Structured Outputs features remain the standard, recommended approach.
Assessing Utility in Production
For organizations handling high-frequency tasks—such as sorting customer inquiries or automated quality control—the Decisions API offers a more efficient path than traditional LLM prompting. By counting monthly input tokens, developers can accurately forecast expenses without the variability introduced by output length. The primary task for adopters in the coming days is to calibrate their confidence thresholds against real-world data to ensure the probabilities returned by gpt-6-luna align with internal accuracy requirements.

