Evidence ledger
Every important statement carries a status, source trail, counterargument, and unresolved question. Frequency is never treated as proof.
Evidence
The public API documentation defines `noul`, `choice`, and `score` response shapes.
Counterargument
A documented interface demonstrates the contract, not the quality of decisions behind it.
Open question
How stable are these contracts and calibration properties across model revisions?
Sources
Evidence
The launch material reports large latency multiples; X discussion mostly repeats those figures.
Counterargument
No independent benchmark in the collected dataset reproduces the headline range on representative workloads.
Open question
What are p50/p95 latency and accuracy under equal task definitions and concurrency?
Evidence
Published pricing is echoed across launch discussion, but remains mutable vendor pricing.
Counterargument
Application cost also includes retries, state construction, integration, and any fallback LLM calls.
Open question
Will pricing and limits remain attractive at production volume?
Evidence
The API contract constrains output types; correctness and calibration are separate empirical properties.
Counterargument
Marketing language such as “zero hallucinations” may use hallucination narrowly to mean invalid free text.
Open question
How should incorrect but schema-valid decisions be measured and communicated?
Sources
Evidence
Those categories recur in the collected launch discussion and align with the documented output primitives.
Counterargument
The sample is launch-week and query-conditioned, so repeated framing does not establish adoption.
Open question
Which use case produces independent, reproducible value first?
Evidence
13 retained posts were classified as code or demos; public repositories exist for Home Assistant, MCP, and a MAGI-style experiment.
Counterargument
Existence of code is not evidence of production reliability or commercial demand.
Open question
Which projects have active users, evaluations, and maintained integrations?
Evidence
Jev produces decisions rather than prose, and Vercel exposes it through an evaluation-oriented API.
Counterargument
Simple rules or conventional classifiers may be cheaper and more predictable for many bounded tasks.
Open question
At what ambiguity and volume does Jev outperform rules, embeddings, and compact classifiers?
Evidence
The product design follows from published pricing and parallel question primitives, not independent production evidence.
Counterargument
Network latency, data preparation, correlated errors, and rate limits may dominate at high decision counts.
Open question
Does batching many questions preserve accuracy and calibration?
Evidence
The collected discussion overwhelmingly relays launch claims; at least one source explicitly labels the figures self-reported.
Counterargument
The ecosystem is only days old, so absence of independent evidence is expected rather than disconfirming.
Open question
Who will publish the first task-matched, reproducible comparison?