Perplexity Decisions API: what it is and what it costs
2026-10-02 · 4 min read
The Perplexity Decisions API is a new endpoint that answers fixed questions with probabilities instead of writing text: yes or no, pick one of these options, or score this against a rubric. It costs $0.04 per million input tokens, output is free, and the model behind it, pplx-decider-v1-27b, is open source under Apache 2.0.
Perplexity's developer account announced it on Thursday. The @perplexitydevs post describes a "multimodal decision model" trained to output "a probability distribution over a fixed set of answers." A few hours later CEO @AravSrinivas confirmed the weights are going open and added that Perplexity intends "to bring down the price even further over the coming days."
How the Decisions API works
You don't write a prompt and parse whatever comes back. Per Perplexity's quickstart, you send a state (text, JSON, or images) plus one or more named questions, and each question has a type:
- A yes/no question gets back the probability that the answer is yes, from 0 to 1.
- A choice question gets back the most likely option, plus a probability for every option.
- A score question takes a rubric of up to 10 ordered levels and returns a probability-weighted average.
A single request can carry up to 128 questions, so one support email could get checked for urgency, department, sentiment and refund risk in one call. Inputs have to stay under about 262,000 tokens, and the rate limit is 10 requests per second per organization. The suggested uses are what you'd guess: sorting content, routing support tickets, grading against a rubric, and pulling structured data out of messy input.
The answer can only be one of the options you defined. A chat model asked to pick "billing, technical or sales" will sometimes reply "this seems like a billing question, but..." and break the code reading it. This one can't.
What it costs
Input is $0.04 per million tokens and output is free. Per decision, that rounds to almost nothing: classifying 10,000 support emails of around 500 tokens each comes to about 20 cents. Images count too. Perplexity says it bills roughly 1,000 input tokens per megapixel.
You can also skip the API completely. The model is on Hugging Face under Apache 2.0, fine-tuned from Qwen3.8-27B. It needs a GPU with roughly 49 GiB of memory just for the weights, though, so self-hosting only makes sense if you already run serious GPU hardware.
How good is it?
The 85.71% figure in the launch post comes from the model card, which averages 11 decision tasks and puts the base Qwen model at 74.76%. A few of those gains are large: RAGTruth, a test of whether an answer is supported by its source text, goes from 61.53% to 88.80%.
What the card doesn't include is a head-to-head against the other decision models. Perplexity is entering a crowded field. @AGTPinsights called it "the next entry in the decision-model wave," arriving the same day as Cloudflare's Clef. The wave started when TypeSafe launched Jev in September, which we covered in a ChatGPT co-inventor built an AI that won't write a word. Jev is closed and text-only. Perplexity's model is open and also reads images. Neither of those facts tells you which one is more accurate on your data.
My take is that the open weights matter more than the 85.71%. A business can test this model, run it privately, and never get stuck if the hosted price or terms change. Arav's promise of further price cuts also suggests Perplexity sees the API as a way to grab developer mindshare more than as a profit line.
What this means for a small business
Most of the useful automation in a small company comes down to small decisions. Is this lead real or spam? Is this invoice a duplicate? Does this review need a reply today? Until now the usual answer was to send those to a big chat model, pay for every word it wrote back, and hope the format held. Decision models make those calls cheap and predictable enough to run on every single message.
Before you trust one, pull a few hundred of your own past decisions, run them through it, and see where it disagrees with what your team actually did. Keep the step that decides separate from the step that writes the reply, so you can swap in whichever decision model wins later on.
If you're not sure which decisions in your week are worth handing off, New Face Design's free process audit maps where your hours go and which calls a model like this could make for you.