Jev:为软件而生的
结构化决策模型
Jev 是 TypeSafe 的旗舰模型,也是第一个 System One 模型。它不生成文本、不写代码、不对话—— 它接收一份 state(状态) 和一组带类型的问题,直接返回带概率分布的 结构化答案,让你的代码可以分支、排序与路由。
01 · Jev 是什么
大语言模型(LLM)被设计成产出给人读的文本。当你的代码实际需要的是模型的一个判断时, 就会产生错配:你在强迫一个文本生成系统输出结构化决策,再把结果解析回代码能依赖的形式。
TypeSafe 的判断是:大规模自动化将由 AI-to-AI、AI-to-Software 的交互主导,机器接口比聊天接口更重要。 因此他们把「读起来舒服」这个目标,换成了「在软件里行为可预测」:具备结构化、可靠性、可观测性、 可测试性、速度、一致性与低成本这些类软件属性。
一次请求,一次响应,全部并行
Jev 的四个关键承诺:
类型安全 by construction
答案永远约束在你给出的选项或等级内。代码不需要从生成的散文里「捞」值,也不存在 JSON 解析失败。
并行且互相独立
一次请求里的每个问题独立并行评估。加问题几乎不改变响应时间,也不会互相污染(无 context-rot)。
校准过的概率
RLCD 校准后的概率跨越大量预测成立:被判 0.8 的这批预测约 80% 真的发生。不确定性成为可用信号。
快到可以放进 UI
多数查询约 100ms 完成,适合实时请求路径、交互界面、游戏,以及大规模离线 Map-Reduce。
Jev 不是聊天 / 代码补全 LLM,不能替代 Claude Code、Cursor、Copilot 等编码智能体背后的模型。 它没有「两种模式」的开关,也不产出解释。正确用法是:继续用你的编码智能体写业务代码, 在需要快速、低成本、结构化判断的地方调用 Jev。
02 · System One 与 RLCD
System One 是一类模型:为软件做快速、结构化决策而生。名字取自 Kahneman《思考,快与慢》—— System 1 是快而直觉的,System 2 是慢而审慎的;这里强调的是「快而聚焦的判断」。
与 LLM 的差异
| 维度 | 通用 LLM | System One(Jev) |
|---|---|---|
| 输出 | 自然语言文本、代码 | 类型化的决策 + 概率分布 + 置信度 |
| 消费方 | 人 | 代码 |
| 答案空间 | 开放 | 由你通过 primitives 约束 |
| 不确定性 | 通常不表达,且趋于过度自信 | 校准概率,可阈值化 |
| 失败模式 | 幻觉、格式错误、啰嗦 | 字面化理解、数学 / 计数弱、不耐多跳间接 |
| 典型延迟 | 数百毫秒到数十秒 | ≈100ms,且加问题几乎不增加延迟 |
三条后训练路线:RLHF / RLVR / RLCD
RLHF
Reinforcement Learning from Human Feedback:把预训练模型变成聊天机器人,优化「人更喜欢」。由 TypeSafe 联合创始人 Diogo Almeida 共同发明。
RLVR
Reinforcement Learning with Verifiable Rewards:造就了擅长数学等可验证任务的推理模型,但更慢、更贵。
RLCD ★
Reinforcement Learning for Calibrated Decisions:TypeSafe 的第三条路——训练模型返回决策与校准概率,而非生成文本。
校准意味着什么
在一个校准良好的模型上,跨大量预测来看:
- 被赋予概率
0.2的结果,大约有 20% 真的发生; - 被赋予概率
0.8的结果,大约有 80% 真的发生; - 被赋予概率
1.0的结果,应当 100% 发生。
这些比率描述的是一组预测,不是对单个答案的保证。
1) 迎合性幻觉:RLHF 奖励「人更喜欢的话」,会连带奖励阿谀与自信口吻的编造。
2) Mode dropping(模式丢弃):偏好优化让模型偏向某种风格(如指令遵循),
压低其它可能输出的概率 —— 这是 GAN 中 mode collapse 的温和版本。
一句极其重要的话:「读起来让人信服」与「可靠到可以无人值守自动化」是两个不同的优化目标。
03 · 三个 AI 原语
TypeSafe 暴露三个 AI 原语,类似软件原语:模块化、可组合、结构化、可靠、快速。每种回答一种形状的问题。
| 原语 | 回答什么 | 适用场景 | 返回字段 |
|---|---|---|---|
| Choice | 是哪一个? | 选项集合固定且无序:路由到部门、文档分类、识别编程语言 | choice、probabilities、confidence |
| Score | 落在哪一档? | 答案在可描述的谱系上:缺陷严重度、客户情绪、技能水平 | score、legend、probabilities、confidence |
| Noul | 这句话是真的吗? | 干净的是 / 否,且概率本身就是有用信号:是否含 PII、是否要求退款 | noul(0–1) |
三种类型可以自由混用于同一个请求。所有问题在同一份 state 上并行、独立评估, 加问题几乎不改变延迟,只增加极便宜的问题 token。 实测(13 个问题合并 vs 13 次单独调用):便宜 12.2 倍、快 10 倍,答案无变化。
原子问题,在代码里组合
把每个问题想成一次「直觉式判定」:一个懂行的人在给到正确上下文后几秒钟内能做的那种判断。 如果一个问题需要长时间推理、或同时权衡多个独立因素,就拆开。
不要问「给这个创业路演打分」。而是分别问:市场规模、技术可行性、差异化程度, 再用你自己的公式组合这些分数。优先级变了,就改代码里的一个系数,而不是重写提示词。
04 · 快速上手
路径 A · Playground(先看效果,不写代码)
- 打开 Playground 并登录;
- 把任意文本粘成 state;
- 加一个 Noul 问题:
"Does this message express urgency?" - 继续加 Choice / Score,混在一次调用里看全部结果。
Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.
路径 B · HTTP API(cURL)
- 从 dashboard 拿 API key;
POST https://api.typesafe.ai/v1/systemone,Header 带Authorization: Bearer <KEY>;- 完整字段见 §20 HTTP API 参考。
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.",
"model": "jev-latest",
"questions": {
"urgency": {
"type": "noul",
"instructions": "Does this message express urgency?"
}
}
}
EOF
完整请求体:三种类型混用
{
"state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.",
"model": "jev-latest",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this",
"criteria": {
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions"
}
},
"frustration": {
"type": "score",
"instructions": "How frustrated the customer appears",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language"
]
},
"is_urgent": {
"type": "noul",
"instructions": "The message conveys urgency or time-sensitivity"
}
}
}
完整响应体
{
"model": "jev-1.13.0",
"answers": {
"department": {
"type": "choice",
"choice": "technical",
"confidence": 0.78,
"probabilities": { "technical": 0.85, "sales": 0.0, "billing": 0.15 }
},
"frustration": {
"type": "score",
"score": 1.0,
"confidence": 1.0,
"legend": {
"0": "Calm, just stating facts",
"1": "Frustrated but civil",
"2": "Very angry, strong language"
},
"probabilities": { "0": 0.0, "1": 1.0, "2": 0.0 }
},
"is_urgent": { "type": "noul", "noul": 1.0 }
},
"usage": { "input_tokens": 392, "output_tokens": 65 }
}
路径 C · Python SDK
pip install typesafe-sdk # 要求 Python >= 3.10
# 或
uv add typesafe-sdk
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
client = TypeSafeClient() # 读取环境变量 TYPESAFE_API_KEY,默认模型 jev-latest
ticket = "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP."
response = client.system_one(
state=ticket,
questions={
"department": Choice(
instructions="Which team should handle this",
criteria={
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions",
},
),
"frustration": Score(
instructions="How frustrated the customer appears",
criteria=[
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language",
],
),
"is_urgent": Noul(
instructions="The message conveys urgency or time-sensitivity",
),
},
)
print(response.answers["department"].choice) # "technical"
print(response.answers["frustration"].score) # 1.0
print(response.answers["is_urgent"].noul) # 1.0
路径 D · 让编码智能体帮你写
# Claude Code
claude plugin marketplace add typesafe-ai/skills
claude plugin install typesafe@typesafe-ai
# 其他智能体(交互式选择)
npx skills add typesafe-ai/skills --skill typesafe-ai # 加 -g 全局安装
装好后给智能体一个起点式提示词:
Let's build a simple CLI that uses the TypeSafe API to evaluate a set of supplied
documents on multiple dimensions. Use the TypeSafe skill to understand how to use the
TypeSafe API and how to structure the system. Ask me questions about what kinds of
documents I want to evaluate and on what dimensions.
05 · State 状态
State 是你请模型审视的内容:一条客服消息、一段文本、或你的应用当前状态。
它填在请求的 state 字段里。一次请求 = 一份 state + N 个问题;所有问题看到同一份 state,彼此独立。
三种合法形状
| 格式 | 适合 | 示例 |
|---|---|---|
| String | 一条消息、一篇文章、一个段落 | "My card was charged twice." |
| Object | 具名字段、相关记录、应用状态 | {"message": "...", "order_id": "A-104"} |
| Array | 一串消息或记录 | ["Hi", "My customer number is TS1337.", "..."] |
大多数请求建议用 Object,让每一部分都有描述性名字,字段之间的关系保持清晰。
{
"ticket": {
"subject": "Duplicate charge",
"messages": [
{"from": "customer", "text": "I was charged twice for order A-104. Please refund the duplicate."},
{"from": "support", "text": "We are checking the charges."}
]
},
"order": {
"id": "A-104",
"charges": [
{"amount_usd": 49, "status": "captured"},
{"amount_usd": 49, "status": "captured"}
]
},
"refund_policy": "Duplicate charges are eligible for a refund."
}
state 里放内容和支持事实(退款请求、政策原文);questions 里放判断 (是否请求退款?政策是否支持?)。把需要互相比较的东西一次性放进同一份 state —— 上面的例子虽然含会话 + 订单 + 策略,但它们是一份 state。
Jev 只接受文本。不支持图片 / 音频 / 视频(截至 jev-1.13)。多模态输入请先在代码里转成文本或结构化字段。
主训练语言是英语;包括 CJK 在内的其它语言可以处理但准确率较低,上线前务必用自己的语料测试,
并重点观察 confidence。
06 · Questions 问题
每个问题由四部分组成:ID、type、instructions,
以及(Choice / Score 必需的)criteria。Noul 的 criteria 可选,用于澄清「是」与「否」的边界。
如何选类型
- Choice:答案落在已知集合、且集合无序。把完整列表给它;担心漏项就加
other/none of the above。 - Score:答案落在一个你能逐点描述的谱系上。
- Noul:干净的是 / 否,且概率本身就是有用信号。
两种都像?选答案能直接被代码消费的那个:Choice 的结果直接映射多条代码路径;Score 映射到一个阈值;Noul 映射到一个 if。
「候选人 Python 强吗?」缺少「强」的定义。Noul 返回 0.5 意味着是与不是的概率各半, 不代表候选人水平中等。想测水平请用 Score 定义等级(无经验 / 略有了解 / 日常使用 / 深度专家); 想要是 / 否就写清楚条件(「简历是否写明在工作中使用过 Python?」)。
用反引号路径引用 state 内部字段
当 state 是 JSON 对象时,可以用带反引号的点 + 索引路径指定要看哪一部分:
questions = {
"refund_requested": {
"type": "noul",
"instructions": "Does `ticket.messages[0].text` request a refund?",
},
"policy_supports_refund": {
"type": "noul",
"instructions": (
"Does `refund_policy` support the refund requested "
"in `ticket.messages[0].text`, given `order.charges`?"
),
},
}
一次打包多个问题
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
state = {
"ticket_message": "My flight was cancelled. Can I get a refund?",
"refund_policy": "Cancelled flights are eligible for a full refund.",
}
with TypeSafeClient() as client:
response = client.system_one(
state=state,
questions={
"refund_requested": Noul(
instructions="Does `ticket_message` request a refund?",
),
"request_type": Choice(
instructions="What is the main request in `ticket_message`?",
criteria={
"refund": "The customer wants money returned.",
"rebooking": "The customer wants a replacement flight.",
"information": "The customer is asking for information only.",
},
),
"frustration": Score(
instructions="How frustrated does the customer appear in `ticket_message`?",
criteria=[
"Calm and neutral.",
"Concerned but civil.",
"Very angry or using strong language.",
],
),
},
)
print(response.answers["refund_requested"].noul)
print(response.answers["request_type"].choice)
print(response.answers["frustration"].score)
什么时候才需要发第二次请求
同一请求里的问题是独立的:一个答案不会成为另一个问题的上下文。发第二次请求的前提是 你的代码必须拿到第一个答案才能构造第二次请求的 state 或选项:
- Skill suggestion:先一次请求给 182 个技能排序,再取前三的完整文本重新评判;
- Structure recovery:先判断每个换行是否切断了句子,据此合并成「块」,再分类这些直到此刻才存在的块;
- Hierarchical classification:每个 Choice 的答案决定下一层提供哪些选项。
除此之外,请合并到一次请求里,让代码忽略它不需要的答案。
07 · Answers 与校准概率
| 类型 | 答案字段 | 怎么读 |
|---|---|---|
| Choice | choice、probabilities、confidence |
choice 是被选中的选项;probabilities 是全部选项上的分布;confidence 概括分布有多尖锐。 |
| Score | score、legend、probabilities、confidence |
score 是沿你等级的位置,可以落在两档之间;legend 把等级编号映射回描述。 |
| Noul | noul |
为「是」的概率。接近 1 强是,接近 0 强否,接近 0.5 不确定。没有独立的 confidence。 |
答案之所以可组合,靠两条性质:
- 答案被约束在你提供的选项内。模型只在你给的选项或等级上返回分布,不会凭空造值。
- 每个答案彼此独立。一个问题的答案不会成为另一个问题的隐藏上下文,增删问题不影响其它答案。
为什么 Noul 没有 confidence
Noul 的分布只有「是 / 否」两个结果,单个 noul 值已经把分布完全描述出来了。
Choice / Score 把概率摊在多个选项或等级上,才需要 confidence 来概括这个摊开的形状。
score = 1.0 可能是「全部概率压在等级 1」,也可能是「等级 0 与等级 2 各占一半」——
两种情况的实际含义天差地别。因此做落地决策时要连读 probabilities 与 confidence。
08 · Confidence 置信度
所有 Choice 与 Score 答案都带 probabilities。分布的形状告诉你模型有多确定;
confidence 是把这个形状压扁成 0–1 的单一统计量,让你不必自己做数学。(Noul 没有这个字段。)
confidence 由 probabilities 推导
TypeSafe 帮你算好并返回。分布越平 → 置信度越低;单点尖峰 → 置信度高。
示意:同样的候选数量下,概率集中度不同 → confidence 差别巨大。
confidence 是「大多数场景合适」的便利度量。你也可以用自己的统计量替代 ——
这正是 TypeSafe 在响应里始终给出完整 probabilities 的原因。
「我不知道」是有用的信号
一个有智能的系统,如果不能诚实表达不确定性,就不能被信任。confidence 给了代码一套分档行为的基础。
代码里的三条路径
| 档位 | 行为 |
|---|---|
| 高置信度 | 自动执行。不需要人介入。 |
| 中置信度 | 谨慎推进:请用户确认、打标复核、或再取一点信息再动。 |
| 低置信度 | 不行动:转人工、请求澄清、或退回另一套系统。 |
阈值随风险缩放
一个置信度阈值从来不是「一个数字」。同一系统内不同动作应按后果分别设门:
response = client.system_one(
state=user_message,
questions={
"action": Choice(
instructions="What is the user trying to do?",
criteria={
"check_balance": "View account balance",
"approve_transfer": "Approve the pending withdrawal request",
"support": "Get help with an issue",
},
),
},
)
action = response.answers["action"]
confidence = action.confidence
if confidence < 0.5:
# 模型真的不确定,不要猜
route_to_human(user_message)
elif action.choice == "check_balance":
# 低风险:显示错页面是可以恢复的
show_balance(account_id)
elif action.choice == "approve_transfer":
if confidence > 0.9:
# 高风险 + 高置信度:确认后执行
confirm_then_execute(account_id)
else:
# 高风险 + 中等置信度:先验证
ask_user_to_confirm(account_id)
正确的阈值取决于你的领域和模型在你数据上的实际表现。从保守阈值开始,用自己的数据观测后调整。
另外:如果你只关心「选最好的那一个」,直接取概率最高的即可,不必设阈值;
如果你心里有特定统计算法,那应该用 probabilities 而不是 confidence。
09 · Choice 选择
从一组已定义的选项中挑一个。返回值包含被选选项、每个选项的概率,以及置信度。
"What programming language is this code written in"
→ options: python, javascript, typescript, go, rust, other
"What type of meeting is this based on the title and description"
→ options: standup, planning, retrospective, one on one, brainstorm, none of the above
"Which product category does this item belong to"
→ options: electronics, clothing, home garden, food and beverage
请求结构
type:恒为"choice"。instructions:模型要回答的问题。criteria:答案选项 map。key 是选项名,value 是该选项的描述。
一个 Choice 最多 255 个选项。选项名与描述都会发给模型,描述要能把彼此区分开。
不需要额外描述的选项可以写 null。
from typesafe_sdk import Choice, TypeSafeClient
with TypeSafeClient() as client:
response = client.system_one(
state="My running shoes arrived in the wrong size. Can I swap them for a size 10?",
questions={
"department": Choice(
instructions="Which team should handle this?",
criteria={
"returns": "Exchanges, wrong or damaged items",
"shipping": "Delivery status, delays, lost packages",
"billing": "Charges, invoices, payment problems",
},
),
},
)
print(response.answers["department"].choice) # "returns"
响应示例:五个 Choice 一次性返回
场景:一张更模糊的工单 —— 涉及三个团队,且没说客户想要什么。其中两个问题是投机性的。
{
"model": "jev-1.13.0",
"answers": {
"department": {
"type": "choice", "choice": "returns", "confidence": 0.42,
"probabilities": { "shipping": 0.04, "billing": 0.35, "returns": 0.61 }
},
"return_reason": {
"type": "choice", "choice": "wrong_size", "confidence": 1.0,
"probabilities": { "other": 0.0, "wrong_size": 1.0, "changed_mind": 0.0, "damaged": 0.0, "wrong_item": 0.0 }
},
"shipping_issue": {
"type": "choice", "choice": "delayed", "confidence": 0.67,
"probabilities": { "wrong_address": 0.0, "other": 0.26, "not_delivered": 0.0, "damaged_in_transit": 0.0, "delayed": 0.74 }
},
"requested_resolution": {
"type": "choice", "choice": "refund", "confidence": 0.2,
"probabilities": { "replacement": 0.34, "refund": 0.4, "information": 0.02, "exchange": 0.24 }
},
"tone": {
"type": "choice", "choice": "frustrated", "confidence": 0.76,
"probabilities": { "frustrated": 0.84, "angry": 0.16, "calm": 0.0 }
}
},
"usage": { "input_tokens": 589, "output_tokens": 212 }
}
department是returns(0.61),但因为提到了重复扣款,billing拿到 0.35 —— 这张工单同时属于两个团队,置信度 0.42 如实反映了这一点。return_reason是wrong_size,置信度 1.0 —— 工单里说得很明白。shipping_issue分散在delayed与other。作为投机性问题时可以直接被代码忽略。requested_resolution偏向refund(0.40),但replacement和exchange分走了大部分;置信度仅 0.20 —— 客户从没说到底要哪一个。tone是frustrated(0.84),置信度 0.76。
把消费逻辑写成普通 if
def triage(ticket: str) -> None:
with TypeSafeClient() as client:
response = client.system_one(state=ticket, questions=TRIAGE_QUESTIONS)
answers = response.answers
department = answers["department"]
if department.confidence < 0.3:
# 不清楚给哪个团队,让人决定
send_to_manual_triage(ticket)
return
if department.choice == "returns":
# return_reason 的答案只在这里使用
assign(ticket, team="returns", issue=answers["return_reason"].choice)
elif department.choice == "shipping":
assign(ticket, team="shipping", issue=answers["shipping_issue"].choice)
else:
assign(ticket, team="billing")
# 概率占比真实的第二个团队也收到副本
for team, probability in department.probabilities.items():
if team != department.choice and probability > 0.25:
notify(ticket, team=team)
resolution = answers["requested_resolution"]
if resolution.confidence < 0.5:
# 客户没说要什么。去问,不要猜。
ask_customer_what_they_want(ticket)
elif resolution.choice == "refund":
flag_for_refund_approval(ticket)
if answers["tone"].choice == "angry":
flag_for_senior_agent(ticket)
对上面那张工单:这段代码把票分给 returns 团队并注明问题是 wrong_size;
因为 billing 的 0.35 超过 0.25 阈值而给它抄送一份;
因为 resolution 置信度 0.20 低于 0.5 而去问客户想要什么;shipping_issue 的答案被完全忽略。
Choice 最多 255 个选项,每个选项只花几个 token —— 所以把完整的团队表 / 类目表 / 产品表都给它,
而不是给一个短名单。担心覆盖不全就加 other 或 none of the above。
10 · Score 评分
Score 用于把内容对着有序的、描述性的等级打分。答案包括 score、每个等级的概率和置信度。
请求结构
type:恒为"score"。instructions:模型在给什么打分。criteria:有序数组,从量表低端排到高端。至少两级,最多 10 级。
Levels:编号从 0 开始
等级编号就是它在 criteria 数组里的位置,从 0 起。模型只看得到描述,
看不到编号也看不到邻居 —— 每个等级都是独立对着 state 判的。
from typesafe_sdk import Score, TypeSafeClient
with TypeSafeClient() as client:
response = client.system_one(
state="The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari.",
questions={
"bug_severity": Score(
instructions="How severe is the reported issue?",
criteria=[
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists",
],
),
},
)
print(response.answers["bug_severity"].score) # 1.43
响应:score 是概率加权均值
{
"model": "jev-1.13.0",
"answers": {
"bug_severity": {
"type": "score",
"score": 1.43,
"confidence": 0.35,
"legend": {
"0": "Cosmetic; no impact to functionality",
"1": "Broken or degraded feature, but workaround exists",
"2": "Blocking issue; no workaround exists"
},
"probabilities": { "0": 0.0, "1": 0.57, "2": 0.43 }
}
},
"usage": { "input_tokens": 332, "output_tokens": 18 }
}
score = 0×0.0 + 1×0.57 + 2×0.43 = 1.43,意为模型在等级 1 与 2 之间分裂、略偏向 1,
置信度 0.35 也是因为这个分裂。注意:这不代表「43% 的客户没有变通方法」 ——
分数是等级编号的加权均值,不是数量比例。
不同输入的对比
| State | score | confidence | L0 | L1 | L2 |
|---|---|---|---|---|---|
| 导出按钮在设置页歪了几个像素。 | 0.0 | 1.0 | 1.0 | 0.0 | 0.0 |
| PDF 导出按钮点了没反应。我还能导出 CSV 再自己转,但太慢了。 | 1.0 | 1.0 | 0.0 | 1.0 | 0.0 |
| PDF 导出失败、转圈不停。团队里有人说 CSV 还能用,有人说也失败了。 | 1.11 | 0.84 | 0.0 | 0.89 | 0.11 |
| Safari 里导出按钮让设置页崩溃。Chrome 正常,但少数客户只用 Safari。 | 1.43 | 0.35 | 0.0 | 0.57 | 0.43 |
| 今天早上起我们全组都登录不了,每次都是 500 错误。 | 2.0 | 1.0 | 0.0 | 0.0 | 1.0 |
注意最后一列的关系:confidence 1.0 只说明「返回的概率分布把全部概率压在一档上」,
不代表答案一定正确。
写好等级的规则
- 描述情境,不要描述程度。「功能损坏但有变通方案」有东西可比;「中等严重」什么都没有。
- 不要写数字。当等级只有
["0","1","2"]时模型没有可比对象 —— 上面那条「歪了几个像素」的报告会从 0.0/1.0 退化成 score 0.55、confidence 0.33。 - 一个 Score 只测一个维度。「准时且聪明且有经验」是三件事,无法放置。
- 极端尾部单独成级。情绪量表最高档若是「非常愤怒」,建议再加一级「辱骂或威胁」。
- 没有中间态就用 Choice;可用 3 级,最多 10 级,别加你描述不清的等级。
拆复杂判断 → 归一化 → 加权
def normalized(answers, question_id: str) -> float:
"""把 score 除以最高等级编号,压到 0–1。"""
top_level = len(TRIAGE_QUESTIONS[question_id].criteria) - 1
return answers[question_id].score / top_level
def priority(ticket: str) -> float:
with TypeSafeClient() as client:
response = client.system_one(state=ticket, questions=TRIAGE_QUESTIONS)
answers = response.answers
severity = normalized(answers, "severity") # 0.62
frustration = normalized(answers, "frustration") # 0.64
report_quality = normalized(answers, "report_quality") # 1.00
# 详细的复现步骤有助于工程师排查,因此略微提升优先级
return 0.6 * severity + 0.3 * frustration + 0.1 * report_quality
# → 0.6 × 0.62 + 0.3 × 0.64 + 0.1 × 1.0 = 0.664
不同量表长度不同:四级量表返回 0–3,三级返回 0–2。不归一化的话,「某量表满分」天然比另一个大, 权重就失去了你以为的含义。
结构化等级 + examples 的实测效果
"criteria": [
{ "what": "Cosmetic; no impact to functionality",
"examples": ["typo in a label", "misaligned icon"] },
{ "what": "Broken or degraded feature, but workaround exists",
"examples": ["export fails in one browser but works in another"] },
{ "what": "Blocking issue; no workaround exists",
"examples": ["cannot log in", "data loss"] }
]
| 等级描述 | score | confidence |
|---|---|---|
| 纯字符串,无 examples | 1.43 | 0.35 |
| 加入相关示例:"export fails in one browser but works in another" | 1.03 | 0.96 |
| 加入无关示例:"search fails, but browsing categories still works" | 1.43 | 0.35 |
示例会引导模型,但只有长得像你真实输入的示例才有用;无关的示例等价于没加。 而且:更高的置信度并不证明答案更正确 —— 请用已知期望等级挑示例,并在另外的输入上验证后再保留。
11 · Noul 是否
Noul 让模型判断一个是 / 否问题,并返回「答案为是」的概率。
请求结构
type:恒为"noul"。instructions:是 / 否问题,或一句让模型判断的陈述。criteria:可选,含true/false两段描述的对象。
from typesafe_sdk import Noul, NoulCriteria, TypeSafeClient
with TypeSafeClient() as client:
response = client.system_one(
model="jev-latest",
state="I have asked three times now. Can I please just talk to a real person?",
questions={
"is_human_escalation": Noul(
instructions="Is the customer asking for a human agent?",
),
"is_repeat_contact": Noul(
instructions="Has the customer contacted support about this before?",
criteria=NoulCriteria(
true="Mentions a prior attempt, ticket, or that they have asked before",
false="No sign of any previous contact",
),
),
},
)
print(response.answers["is_human_escalation"].noul) # 0.99
print(response.answers["is_repeat_contact"].noul) # 0.93
怎么读 Noul(实测 jev-1.13.0)
| State | noul |
|---|---|
| Thanks, that fixed it! | 0.02 |
| How do I reset my password? | 0.07 |
| I need this sorted today, whatever it takes. | 0.26 |
| Are you a bot? | 0.40 |
| Is there any way to speak to someone about my invoice? | 0.84 |
| I have asked three times now. Can I please just talk to a real person? | 0.99 |
「今天必须搞定」很紧急但没有要求人工 → 0.26;「你是机器人吗?」暗示想要人但没明说 → 0.40。 这两类正是需要用代码里的阈值来兜的场景。
Noul 不是程度量尺
| 候选人 | Noul:"候选人 Python 强吗?" | Score:"候选人 Python 经验多少?" |
|---|---|---|
| 我的经验是 Java 和 Go,没用过 Python。 | 0.03 | 0.0(无经验) |
| 偶尔用小脚本,主职是 Java。 | 0.14 | 1.0(略有了解) |
| 上一份工作两年每天用 Python,主要是数据管道。 | 0.81 | 2.05(工作中常规使用) |
| 写了八年 Python,维护过大型 Django 代码库。 | 0.92 | 2.89(深度专家) |
Noul 判断的是一个命题「强」,值是它的可能性。你可以在代码里把 0.3–0.7 划为「有些经验」, 但模型并不会看到这些区间,答案里也没有任何东西是对着它们判的。 中间值可能意味着「中等经验」,也可能意味着「这个案例不清楚」;候选之间的间距也不是你选的。
写好 Noul 的四条
- 一个问题只问一件事。「客户既生气又要求退款吗?」模型要同时判两件事,值就没有意义了。拆成两个 Noul,代码里组合。
- 让「高值 = 是」。「这条消息含个人数据吗?」是对的;「这条消息不含个人数据吗?」会把后来的代码读反。
- 陈述句也一样有效。「客户正在请求退款」,接近 1 即陈述为真。两种写法都用你自己的数据试试。
- 让边界无歧义。「候选人有用过 Python 吗?」里的「有」没有中间地带。边界微妙时加
criteria的 true / false 描述。
阈值三档:自动 / 复核 / 拒答
YES = 0.8
NO = 0.2
def route(message: str) -> None:
with TypeSafeClient() as client:
response = client.system_one(
model="jev-latest", state=message, questions=SUPPORT_QUESTIONS
)
answers = response.answers
wants_human = answers["is_human_escalation"].noul
repeat = answers["is_repeat_contact"].noul
if NO < wants_human < YES or NO < repeat < YES:
# 模型两边都不确定,让人决定
send_to_review(message)
return
priority = "high" if repeat > YES else "normal"
if wants_human > YES:
route_to_agent(message, priority=priority)
else:
route_to_bot(message, priority=priority)
阈值怎么定取决于「错的代价」:0.5 用于是 / 否同样容易行动的场景;
误判「是」代价高时抬高(呼叫值班人员、放款);
漏掉「是」代价高时降低(没标出安全问题)。
复核队列太多就收窄这个区间;放过去的错太多就放宽。
结构化 instructions:一份简历 vs N 条候选库记录
SAME_PERSON = "Is the resume for the same person as `potential_duplicate`?"
def duplicate_questions(candidates: list[dict]) -> dict[str, Noul]:
"""每个候选库记录一个 Noul,问同一个问题,全部打包进一次请求。"""
return {
f"same_as_record_{candidate['id']}": Noul(
instructions={
"potential_duplicate": {
"name": candidate["name"],
"location": candidate["location"],
"last_employer": candidate["last_employer"],
},
"question": SAME_PERSON,
},
)
for candidate in candidates
}
实测响应:same_as_record_18 = 0.74(名字拼写不同但地点与雇主一致)、
_42 = 0.09(同名不同城市不同雇主)、_77 = 0.08(近似名同地点不同雇主)。
在代码里各自阈值化,中间值送人工。
12 · 结构化问题(Advanced)
System One 模型是被训练来懂结构的。instructions、Choice 的选项描述、Score 的等级、Noul 的 criteria 全都接受 JSON 结构。
哪里可以结构化
| 字段 | 适用 | 可接受形状 |
|---|---|---|
instructions | Choice / Score / Noul | string、object、array、null |
criteria 值(选项描述) | Choice | 同上 |
criteria 条目(等级描述) | Score | 同上 |
criteria.true / criteria.false | Noul | 同上 |
什么时候值得结构化
- 为了清楚。当问题有多个部分时,JSON 的 key 自带标签。
- 当问题需要附带数据。schema、分类法、数据库行本来就是 JSON —— 直接整体传入或取相关子字段, 不要序列化拼接成字符串模板。
结构化 Choice 选项:用对比式 rubric 拉清边界
先给每个选项一行描述。当两个相似选项被模型反复混淆时,改成对象:包含「覆盖什么」、 「什么属于邻居选项而非本选项」、以及一些示例输入。
{
"return_topic": {
"type": "choice",
"instructions": "What is this ticket about?",
"criteria": {
"return_policy": {
"what": "Asks how returns work: the window, condition rules, who pays shipping",
"not_for": "A question about the state of an already-opened return",
"examples": ["What is your return window?", "Do I pay return shipping?"]
},
"return_status": {
"what": "Asks about a specific in-flight return: where it is, when refund lands",
"not_for": "General questions about how the policy works",
"examples": ["Where is my return?", "When will my refund arrive?"]
}
}
}
}
结果是 return_status,置信度 1.0。字段名 question、focus、what、
not_for、examples都不是 API 的一部分,也不保留 —— 你随便起,
模型会连同值一起看到,所以用短名字准确标注。
遍历分类树
要分到很深的分类树,就每层一个 Choice,在代码里走树:每一步的选项是当前节点的子节点, 每个选项的值是那棵子树。这样模型在提交之前能看到分支下面有什么 —— 当目标叶子节点的名字从分支名看不出来时,这很关键。
例:一个水瓶商品可能同时落在 Sporting Goods > Cycling > Bike Bottles & Cages 与
Home & Kitchen > Drinkware > Water Bottles 之下。把子树展示出来,模型才能权衡
「强调车用水壶架」与「日常水杯」两侧。probabilities 还告诉你这个分裂是否接近,值不值得两条分支都探。
分支过大时,把值裁剪成「直接子节点 + 一部分叶子样本」。参见 Hierarchical Classification cookbook 里的 beam search。
数组形式的 instructions
"instructions": {
"question": "Does the claimed sender identity conflict with the sending domain?",
"compare": ["ticket.sender.display_name", "ticket.sender.email"],
"focus": "Compare the named organization with the email domain."
}
13 · 模式一:投机扇出(Speculative Fan-Out)
把你的系统可能需要的所有问题都塞进一次调用,包括投机性的那些,让代码事后决定哪些相关。
category = response.answers["category"]
bug_severity = response.answers["bug_severity"]
bug_repro = response.answers["has_reproducible_steps"]
refund = response.answers["refund_requested"]
frustration = response.answers["frustration"]
if category.choice == "bug_report":
if bug_severity.score > 1.5 and bug_repro.noul > 0.6:
escalate_to_engineering(ticket_id, severity="high")
else:
add_to_bug_backlog(ticket_id)
elif category.choice == "billing":
if refund.noul > 0.7:
route_to_billing_with_flag(ticket_id, refund_likely=True)
else:
route_to_billing(ticket_id)
elif category.choice == "feature_request":
log_feature_request(ticket_id)
# Frustration 与分类无关,始终有用
if frustration.score > 1.5:
flag_for_priority_response(ticket_id)
完整决策树所需的一切来自一次调用。投机问题在无关时被忽略、在相关时省下一次往返。 13 问合并 vs 13 次单独调用的实测:成本 ↓12.2×、延迟 ↓10×, 且重复运行间的标准差没有变化(多数答案 5 次重复完全一致)。
14 · 模式二:置信度门控路由
把置信度当作第二个决策轴:答案告诉你「是什么」,置信度告诉你「要不要动」。
场景:语音银行指令
同一个系统里,不同操作的后果不同,就该配不同的门限。
action = response.answers["intent"]
# 任何动作的置信度低于 0.6 → 转人工
if action.confidence < 0.6:
route_to_support_agent(account_id)
elif action.choice == "check_balance":
# 低风险:0.6 足够
show_balance(account_id)
elif action.choice == "approve_transfer":
if action.confidence > 0.85:
# 高风险但高置信度:自动执行
approve_transfer(account_id)
else:
# 高风险 + 中等置信度:先和用户确认
ask_user_to_confirm("Just to confirm: you would like to approve this transfer, is that correct?")
else:
route_to_support_agent(account_id)
0.6 这个地板捞住模型真心不确定的一切。地板之上,每个动作按自己出错的代价设阈值:
查余额 0.6 就够(最坏也就是用户听了一遍余额),转账要 >0.85,否则请用户确认。
它们是教学用的数字,不是推荐值。请用你自己的数据画「置信度 vs 准确率」曲线来标定真值。
15 · 模式三:复合评分
把复杂判断拆成互相独立的若干维度,各自打分,再用你控制的权重在代码里合成。
场景:简历筛选
一次请求四个 Score:Python 深度、团队领导、系统设计、通才程度。全部归一化到 0–1 后按岗位加权。
py = response.answers["python_depth"].score / 4
lead = response.answers["team_leadership"].score / 4
arch = response.answers["system_design"].score / 4
general = response.answers["generalist"].score / 4
# Senior IC 权重
ic_score = (0.40 * py) + (0.10 * lead) + (0.40 * arch) + (0.10 * general)
# Engineering Manager 权重
em_score = (0.15 * py) + (0.40 * lead) + (0.20 * arch) + (0.25 * general)
你能清楚看到最终分数是怎么算出来的。排名不符合预期时,改系数再跑一遍, 而不是回去重写一段提示词。收益:成本↓、可靠性↑、速度↑。
16 · 模式四:意图路由
不是每个请求都需要同一种处理器。有的查数据库就行,有的需要带领域上下文的专家 LLM, 有的必须给人。TypeSafe 可以站在最前面,当作一个又快又便宜的分类器来决定调谁。
场景:客服路由
def route_ticket(ticket_id, response):
intent = response.answers["intent"]
complexity = response.answers["complexity"]
if intent.confidence < 0.5:
# 连分类都没有足够把握 → 转人工
return route_to_human_agent(ticket_id)
if intent.choice == "order_status":
handle_order_status(ticket_id) # 确定性代码,不碰 LLM
elif intent.choice == "product_question":
handle_with_llm(ticket_id, PRODUCT_SPECIALIST)
elif intent.choice == "return_exchange":
handle_with_llm(ticket_id, RETURNS_SPECIALIST)
elif intent.choice == "complaint":
low_confidence = complexity.confidence < 0.5
# complexity.score 越高越倾向于「需要升级」
if complexity.score > 1 or low_confidence:
# 太复杂不适合安全自动化,或我们对复杂度也没把握 → 给人
route_to_human_agent(ticket_id)
else:
handle_with_llm(ticket_id, COMPLAINT_RESOLUTION)
一个意图直接落到确定性代码(完全不用 LLM),两个落到不同上下文的专家 LLM, 一个用复杂度分数在「LLM vs 人」之间选择。昂贵的资源只在真正需要的请求上被调用。
低置信度分数在你这个系统语境下的含义、以及这个决策的赌注,永远要一并考虑。
17 · 如何用 TypeSafe 构建系统
一句话总结:先写一个正常的软件工作流,然后只在真正需要 AI 的地方插入 System One。
① 控制流、确定性规则、副作用留在代码里;
② 把宽泛的判断拆成窄而类型化的问题,配显式 instructions 与 criteria;
③ 每个问题只给它需要的上下文;
④ 用概率与置信度决定「行动 / 复核 / 升级」;
⑤ 独立问题一起问,答案在代码里组合。
三种软件架构
| 架构 | 特征 |
|---|---|
| 传统软件 | 由简单软件原语构成的复杂决策树。每个原语可靠,开发者才能把它们组合成更高层抽象。 |
| LLM 智能体 | 处理指令并自行选择下一步。有人盯着时很好用,但每个循环都多一次脱轨机会。 |
| AI 驱动的软件 | 代码做确定性工作、拥有控制流;模型只出现在需要「可编程常识」或解读非结构化数据的地方,且每个 AI 任务都保持原子与受约束。 |
System One 为什么可组合
Structured
类型安全 by construction,结果符合你的结构化类型 / JSON schema。
Parallel
独立并行评估,一个结果不会变成改变别人的隐藏上下文。
Comparable
输出可排序,可驱动 if、阈值与比较。
Calibrated
RLCD 用校准概率表达不确定性,而非趋于过度自信。
设计流程八步
- 能用代码就用代码
确定性工作可靠且便宜。别用智能体的
while循环表达一个软件工作流能说完的行为。pythondays_overdue = (today - invoice.due_date).days if days_overdue > 30: route_to_collections(invoice) - 裁剪输入 state
只带当前问题相关的上下文。不要用模型权重里的知识替代你自己的知识库。
- 在 state 里用结构
嵌套 JSON + 反引号点索引路径(如
`support.tickets[0].message`)指向具体值。 - 拆解问题(最重要的一步)
宽泛的问题把好几个判断藏在一个答案背后;原子问题把这些判断暴露出来,让你能检查、调参、组合。
- 在 question 里用结构
短而无歧义的问题保持字符串;需要上下文 / 示例、或部分由代码生成时才用对象。
- 一次请求里大量提问
问题并行执行,加问题几乎不增加延迟 —— 这是把「每美元智能」最大化的核心手段。
- 在代码里组合输出(或喂给经典 ML 模型)
python
quality = ( 0.4 * answers["answers_request"].noul + 0.4 * answers["citations_are_supported"].noul + 0.2 * (1 - answers["contradicts_context"].noul) )没有标签给下游模型时,可以用一组昂贵的推理模型生成标签(见 AutoResearch cookbook)。
- 按不确定性路由
python
answer = response.answers["card_help_topic"] if answer.confidence < 0.8: route_to_human_review(ticket) else: route_to_handler(answer.choice, ticket)阈值请用「置信度 vs 准确率」曲线在自己数据上标定。
端到端示例:客服工单分流
下面这段把以上八步串起来:确定性前置 → 结构化 state → 原子 + 结构化问题 → 加权组合 → 置信度门控 → 投机答案按需使用。
from typesafe_sdk import Choice, Noul, NoulCriteria, Score, TypeSafeClient
def triage_ticket(ticket, customer):
# 1) 确定性状态不调用模型
if ticket["status"] == "closed":
return "no_action"
open_orders = [
order for order in customer["orders"] if order["status"] != "delivered"
]
# 2) 只带上问题需要的结构化上下文
state = {
"ticket": {"message": ticket["message"], "sender": ticket["sender"], "links": ticket["links"]},
"customer": {"plan": customer["plan"], "open_orders": open_orders},
"policy": {"sensitive_credentials": ["password", "security code", "API key"]},
}
# 3) 结构化 + 原子问题,一次性并行
questions = {
"topic": Choice(
instructions={
"question": "Which team should handle `ticket.message`?",
"focus": "Classify the customer's primary request.",
},
criteria={
"billing": {
"what": "Charges, invoices, refunds, or subscriptions",
"not_for": "Order tracking or account access",
"examples": ["I was charged twice", "Where is my refund?"],
},
"orders": {
"what": "Order status, delivery, cancellation, or returns",
"not_for": "Charges or account access",
"examples": ["Where is my order?", "Cancel my shipment"],
},
"account": {
"what": "Login, profile, permissions, or security",
"not_for": "Charges or order tracking",
"examples": ["Reset my password", "I cannot sign in"],
},
},
),
"requests_credentials": Noul(
instructions={
"question": "Does the message request a sensitive credential?",
"compare": ["`ticket.message`", "`policy.sensitive_credentials`"],
"focus": "Look for a request to disclose the credential itself.",
},
criteria=NoulCriteria(
true={"what": "Asks the recipient to disclose a listed credential",
"examples": ["Reply with your password", "Send us your API key"]},
false={"what": "Does not ask the recipient to disclose a credential",
"not_for": "A legitimate instruction to reset a credential",
"examples": ["Use this link to reset your password"]},
),
),
"sender_identity_mismatch": Noul(
instructions={
"question": "Does the claimed sender identity conflict with its domain?",
"compare": ["`ticket.sender.display_name`", "`ticket.sender.email`"],
"focus": "Compare the named organization with the email domain.",
},
criteria=NoulCriteria(
true={"what": "Claims an organization unrelated to the email domain",
"examples": ["Acme Payroll sent from claim-bonus.example"]},
false={"what": "The identity and domain agree or make no conflicting claim",
"examples": ["Acme Payroll sent from acme.example"]},
),
),
"unexpected_reward": Noul(
instructions={
"question": "Does the message announce an unexpected reward?",
"inspect": "`ticket.message`",
"focus": "Look for an unsolicited prize, payment, or reward claim.",
},
criteria=NoulCriteria(
true={"what": "Announces an unrequested prize, payment, or reward",
"examples": ["You were selected for a $1,000 bonus"]},
false={"what": "Contains no reward claim or discusses an expected payment",
"not_for": "A customer asking about a known refund or payroll deposit",
"examples": ["When will my approved refund arrive?"]},
),
),
"frustration": Score(
instructions={
"question": "How frustrated does the customer appear?",
"inspect": "`ticket.message`",
"focus": "Judge expressed frustration, not issue severity.",
},
criteria=[
{"what": "Calm and matter-of-fact",
"signals": ["Neutral wording", "No complaint about the experience"]},
{"what": "Frustrated but civil",
"signals": ["Expresses annoyance", "Remains constructive"]},
{"what": "Very angry or threatening to leave",
"signals": ["Hostile language", "Threatens cancellation or churn"]},
],
),
}
with TypeSafeClient() as client:
response = client.system_one(state=state, questions=questions)
answers = response.answers
# 4) 加权组合 —— 权重由代码掌控
spam_risk = (
0.45 * answers["requests_credentials"].noul
+ 0.30 * answers["sender_identity_mismatch"].noul
+ 0.25 * answers["unexpected_reward"].noul
)
# 5) 不确定的判断升级,而不是硬猜
spam_is_uncertain = 0.4 < spam_risk < 0.6
if spam_is_uncertain or answers["topic"].confidence < 0.75:
return route_to_human_review(ticket)
if spam_risk >= 0.6:
return quarantine_as_spam(ticket)
# 6) 投机答案只在对应路径上使用
if answers["topic"].choice == "billing":
return route_to_billing(ticket, refund_requested=answers["refund_requested"].noul >= 0.7)
if answers["topic"].choice == "orders":
return route_to_orders(ticket, mentions_open_order=answers["mentions_open_order"].noul >= 0.7)
priority = (
"high"
if answers["frustration"].confidence >= 0.7 and answers["frustration"].score >= 1.5
else "normal"
)
return route_to_account_support(ticket, priority=priority)
18 · Jev 1.13 的棱角(已知失效模式)
Jev 并不完美。jev-1.13 快、校准良好、常识判断不错,但它会在需要额外一层间接的地方绊倒,
理解有时会过于字面,数学方面吃力。以下是官方公开的九个失效模式(最后审阅:2026-09-17)。
| # | 失效模式 | 应当怎么做 |
|---|---|---|
| 1 | 字面化理解 | 把条件写死,每个选项都写清 criteria |
| 2 | 数学与数字 | 算术留在代码里 |
| 3 | 日期与时间比较 | 抽出组成部分,在代码里比较 |
| 4 | 间接指代 | 减少跳数,直接点名 state 的相关部分 |
| 5 | 塞满无关细节的大 state | 先过滤,只发问题需要的 |
| 6 | 对抗性内容 | 提示要精确,上线前测边界用例 |
| 7 | 指令与 criteria 自相矛盾 | 把 criteria 视为指令的延伸并保持一致 |
| 8 | 常识性结构不变量 | 每个判断只用一种方式提问,恒等式在代码里强制 |
| 9 | 生成文本 | 用生成式模型 |
1 · 字面化理解
Jev 回答的是你写下的问题,不是你想问的问题。范围词、否定、隐含条件都会被按字面读。
替代做法:把确切条件写进 instructions,边界用例放进 criteria。
当你看着一个错答案、并开始解释「我其实想问的是……」时,这段解释就是指令里缺失的那一半。
如果确实不可避免要解释,就拆成两个字面问题,在代码里组合。
2 · 数学与数字
计数不可靠。包括单词里的字母数、段落里某词的出现次数、长列表里的条目数。 模型是在「认出答案的形状」,而不是在计数,误差随被计对象的规模增长。
from typesafe_sdk import Noul, TypeSafeClient
client = TypeSafeClient(model="jev-1.13")
YES = 0.5 # 阈值视你的场景而定
items = ["typesafe", "apple", "california", "banana", "likes", "calibration", "orange", "vertex"]
result = client.system_one(
{"items": items},
{
f"item_{i}": Noul(instructions=f"Is `items[{i}]` the name of a fruit?")
for i in range(len(items))
},
)
count = sum(result.nouls[f"item_{i}"].noul > YES for i in range(len(items)))
数值表示也很弱。用十六进制值问颜色会明显劣于英文色名;给 RGB / hex 无法可靠判断两个值是否接近。
高层编程语言的题优于汇编或二进制 —— 语义表达优于数值表达。
不要用 Score 做算术插值:jev-1.13 的等级在数值校准上是弱的,
无法通过在相邻两级之间插值来还原精确数值;但可以用期望值去卡阈值。
3 · 日期与时间比较
Jev 把日期读成文本,不是有序量。问「哪个日期在前」「相隔多久」「是否落在窗口内」都不可靠, 混合格式、相对表述、季度 / 结算窗口等域边界会让它更糟。
替代做法:切开做。抽取是判断 → 给模型;算术不是 → 留在代码。 日期的每个部分都是小闭集(12 个月、31 个可能的日、有限年份), 于是抽取可以变成一个枚举上的 Choice,还能显式放一个「未说明」选项让缺失被报告而非被猜。 代码负责把各部分组装成真日期,并拥有之后的一切:排序、时长、偏移、星期。
4 · 间接指代
含双重否定或复杂间接的指令可靠性更差。「某个属性的属性」、需要多跳推理的问题都会掉点。
替代做法:尽可能直写;能的话把 state 的相关部分按名字点出来。
5 · 塞满无关细节的大 state
state 里与决策无关的内容越多,准确率越低:无关细节是干扰项, 而且大 state 更难定位是哪个部分造成了错答案。
替代做法:先在代码里检索与过滤,只发送问题需要的字段。无法过滤时,
可以用一个 Noul 筛相关性(见 Classifying RAG passages cookbook)。
jev-1.13 上下文窗口有界:整请求 64k tokens;state + 最长那一个问题 为 32k tokens。
6 · 对抗性内容
state 是数据,Jev 默认不把它当敌对的。被写成要刻意引导模型的内容 (注入的指令、刻意误导的框架、论证自身应如何被分类的文本)都可能挪动答案。官方表示未来会改进。
替代做法:criteria 写得明确,集成彻底测试后再放给用户。
7 · 指令与 criteria 自相矛盾
当 instructions 与 criteria 要的是不同的东西,模型可能糊涂。
例如一个 Noul 里 true 映射到「否」、false 映射到「是」,表现就更差。
最佳表现来自清楚的措辞 —— 目标是让普通人也容易读懂。
替代做法:把 criteria 当作指令的延伸,用清晰准确的语言对齐二者。
8 · 常识性结构不变量(重要)
Jev 极其一致 —— 语义相近的输入会得到量化相近的输出。 但你能想到的很多结构恒等式并不被模型保证。
同一个命题同时用 Noul 和 yes / no Choice 提问(针对工单 "I'm not happy with the fit. What are my options here?"):
Noul noul | Choice yes | Choice no | Choice confidence |
|---|---|---|---|
| 0.22 | 0.01 | 0.99 | 0.97 |
命题与其否定分别作为两个 Noul(针对另一条工单):
refund | not_refund | 求和 |
|---|---|---|
| 0.72 | 0.47 | 1.19 |
替代做法:不要依赖预期的结构恒等式。也不要把在 Noul 上调好的阈值原样搬到 Choice 上。 一个多选项的 Choice 和「一个选项一个 Noul」回答的是不同的问题: Choice 是相对的(决定哪一个),每个 Noul 是绝对的(可能全部都低)。 Skill suggestion cookbook 同时使用两者来处理同一份短名单 —— Choice 挑技能,Noul 决定要不要给建议。
9 · 生成
Jev 没被训练来生成文本。虽然可以通过串联 Choice 强行让它生成,但效果差且非常慢。 数据抽取更好的做法是用正则或生成模型先找出候选,让 Jev 挑正确的一个。
✗ 让模型算代码能精确算出来的东西
✗ 把多个判断藏进一个问题里
✗ System Two 任务(多层间接推理)
✗ 给 state 超出问题需要的上下文 —— Jev 会 context rot,无关材料直接损耗准确率
19 · 模型 · 定价 · 限额
所有模型都由同一个端点 POST /v1/systemone 提供服务,请求里的 model 字段决定谁来回答。
- Price:按输入 token 计费,输出 token 免费。Btok = 十亿 token,Mtok = 百万 token。
- Rate limits:以 token/秒 与 请求/分钟 计。超任一限制返回
429 Too Many Requests。SDK 默认带退避重试,并遵守retry-after;直接调 HTTP 请自行实现。 - Context:Jev 只摄入 state 一次,并让所有问题在其上并行评估。64k 预算覆盖 state + 全部问题;32k 预算适用于 state + 单条最长的问题。
- Input:先把非文本输入(图像、音频、视频、二进制)预处理成文本或结构化字段。
官方明示当前需求量大,以上限额可能无通知地变化;定制与企业计划可提供更高上限。
别名
| 别名 | 指向 | 含义 |
|---|---|---|
jev-latest | jev-1.13.0 | 最新的稳定正式发布。SDK 默认值,也是文档示例使用的名字。 |
jev-preview | jev-1.13.0 | 最新发布(无论是否正式)。有预览构建时会跑到 jev-latest 之前。当前两者同指,暂无预览构建。 |
新发布上线时别名会挪,其背后的答案可能在你没改动的情况下变化。
响应的 model 字段会报告实际作答的版本化 ID,便于日志记录。
如果你针对某个特定版本调好了置信度阈值,请钉住版本 ID 而不是别名,按自己的节奏升级。
没有微调,没有 LoRA
Jev 不使用客户数据微调或 LoRA 适配,用同一套权重服务所有账户。你通过请求而非私有权重来塑造它:
- 把专有内容、记录、参考材料放进
state; - 把领域规则与边界用例写进每个问题的
instructions与criteria; - 把宽泛判断拆成原子问题,在代码里组合(含用 Jev 的概率训练下游经典模型)。
其它
- 语言支持:英语为主训练语言且当前最优;含 CJK 在内的其它语言可处理但相对较弱。
- 数据:不使用客户请求与响应进行训练;企业提供 ZDR(零数据留存)。
列出可用模型
GET /v1/models 返回你的账户可以发送的 model 名字(当前列出别名),含描述与发布日期。版本化 ID 无论是否在列表里都可被接受。
curl https://api.typesafe.ai/v1/models \
-H "Authorization: Bearer $TYPESAFE_API_KEY"
from typesafe_sdk import TypeSafeClient
with TypeSafeClient() as client:
for model in client.models.list().models:
print(model.name, model.release_date, model.description)
import { TypeSafeClient } from "@typesafe-ai/sdk";
const client = new TypeSafeClient();
const models = await client.models.list();
for (const model of models) {
console.log(model.name, model.release_date, model.description);
}
响应字段:models[] → name(可被 model 字段接受的 ID 或别名)、description、release_date。
20 · HTTP API 参考
POST https://api.typesafe.ai/v1/systemone
Authorization: Bearer <API_KEY>
Content-Type: application/json
请求体顶层字段
"jev-latest"。map 条目:你选的 key —— 不会发给底层模型,也不参与推理。
Noul question
true:值为「是」(接近 1)意味着什么;false:值为「否」意味着什么。二者均为 string | object | array | null。{
"state": "Help! My payouts have been failing for 3 days.",
"model": "jev-latest",
"questions": {
"is_urgent": {
"type": "noul",
"instructions": "Does this convey urgency?",
"criteria": {
"true": "Explicitly time-sensitive",
"false": "No urgency expressed"
}
}
}
}
Choice question
null。每个 Choice 最多 255 个选项。{
"state": "Help! My payouts have been failing for 3 days.",
"model": "jev-latest",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoicing, refunds",
"technical": "Bugs, outages, integrations",
"sales": "Pricing, upgrades, new accounts"
}
}
}
}
Score question
{
"state": "Help! My payouts have been failing for 3 days.",
"model": "jev-latest",
"questions": {
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Frustrated", "Very angry"]
}
}
}
结构化 instructions(任意问题类型通用)
"instructions": {
"potential_duplicate": {
"name": "John Smith",
"location": "Oakland, California",
"last_employer": "Google"
},
"question": "Is the resume for the same person as `potential_duplicate`?"
}
响应体
input_tokens(integer)、output_tokens(integer)。Answer 类型
{
"model": "jev-1.13.0",
"answers": {
"frustration": {
"type": "score",
"score": 1.05,
"legend": { "0": "Calm", "1": "Frustrated", "2": "Very angry" },
"probabilities": { "0": 0.0, "1": 0.95, "2": 0.05 },
"confidence": 0.92
}
},
"usage": { "input_tokens": 304, "output_tokens": 18 }
}
21 · Python SDK
uv add typesafe-sdk
# 或
pip install typesafe-sdk
要求 Python >= 3.10。客户端自动读取环境变量 TYPESAFE_API_KEY,默认模型 jev-latest。
同步 / 异步
from typesafe_sdk import AsyncTypeSafeClient, Choice, Noul, Score
async def main() -> None:
async with AsyncTypeSafeClient() as client:
response = await client.system_one(
state={"document": "I was charged twice. Please fix this ASAP."},
questions={
"billing": Noul(instructions="Is this ticket about billing?"),
"tone": Choice(
instructions="What is the customer's tone?",
criteria={"calm": None, "frustrated": None, "angry": None},
),
"urgency": Score(
instructions="How urgent is this ticket?",
criteria=["can wait", "this week", "today"],
),
},
)
print(response.nouls["billing"].noul)
print(response.choices["tone"].choice)
print(response.scores["urgency"].score)
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
with TypeSafeClient() as client:
response = client.system_one(
state={"document": "I was charged twice. Please fix this ASAP."},
questions={
"billing": Noul(instructions="Is this ticket about billing?"),
"tone": Choice(
instructions="What is the customer's tone?",
criteria={"calm": None, "frustrated": None, "angry": None},
),
"urgency": Score(
instructions="How urgent is this ticket?",
criteria=["can wait", "this week", "today"],
),
},
)
print(response.nouls["billing"].noul)
print(response.choices["tone"].choice)
print(response.scores["urgency"].score)
核心类型
| 类型 | 字段 / 说明 |
|---|---|
JSONValue | JSON 风格的值,可嵌套、可含 None。 |
JSONContent | 纯字符串,或 JSONValue 组成的 mapping / sequence。 |
NoulCriteria | true、false(各为 JSONContent 或 None);additionalProperties: false。 |
Noul | type="noul"、instructions、criteria(可选)。 |
Choice | type="choice"、instructions、criteria(选项 map)。 |
Score | type="score"、instructions、criteria(有序等级数组)。 |
ChoiceAnswer | 必填:choice、confidence、probabilities。 |
ScoreAnswer | 必填:score、confidence、legend、probabilities。 |
NoulAnswer | 必填:noul。 |
SystemOneResponse | 答案按类型分组成 nouls / choices / scores,另有 answers(按原 ID)、model、usage。 |
Python SDK 的 probabilities 与 legend 按整数等级作为键,而 HTTP 响应里是字符串键。
常量
| 常量 | 值 | 说明 |
|---|---|---|
API_KEY_ENV | TYPESAFE_API_KEY | API key 环境变量 |
BASE_URL_ENV | TYPESAFE_BASE_URL | API 根地址环境变量 |
DEFAULT_MODEL_ENV | TYPESAFE_DEFAULT_MODEL | 默认模型环境变量 |
LOG_LEVEL_ENV | TYPESAFE_LOG_LEVEL | 日志级别环境变量 |
DEFAULT_BASE_URL | https://api.typesafe.ai | 默认 API 根地址 |
DEFAULT_MODEL | jev-latest | 默认模型名 |
DEFAULT_TIMEOUT | 10.0 | 每个 HTTP 操作的默认超时(秒) |
重试策略
from typesafe_sdk import RetryPolicy, TypeSafeClient
client = TypeSafeClient(
retry=RetryPolicy(
max_retries=3, timeout=10.0, http_statuses={429, 500, 502, 503, 504}
)
)
| 参数 | 说明 |
|---|---|
max_retries | 初始尝试之后的最大重试次数;0 关闭重试。 |
backoff_initial / backoff_max | 首次退避秒数、退避上限秒数(每次翻倍);0 关闭退避。 |
backoff_jitter | 每次退避随机减去的比例,0–1。 |
http_statuses | 被重试的 HTTP 状态码集合。 |
respect_retry_after | 是否遵守 Retry-After / retry-after-ms 响应头。 |
api_connection_error | 是否重试 TypeSafeAPIConnectionError。 |
api_timeout_error | 是否重试 TypeSafeAPITimeoutError。 |
exceptions | 在内置规则之外额外触发重试的异常类型。 |
predicate | 可选的判定函数;收到异常时返回 True 即在其它规则之外额外触发重试。 |
timeout | 每次 SDK 调用的总重试预算(秒),含首次尝试与延迟;None 关闭上限。会在某次退避达到或超过预算前停止,并重新抛出最后一个错误。 |
22 · JavaScript / TypeScript SDK
npm install @typesafe-ai/sdk
import { choice, TypeSafeClient } from "@typesafe-ai/sdk";
const client = new TypeSafeClient();
const response = await client.systemOne({
state: { document: "I was charged twice. Please fix this ASAP." },
questions: {
category: choice("What is this ticket about?", {
billing: null,
technical: null,
other: null,
}),
},
});
console.log(response.answers.category.choice);
设置环境变量 TYPESAFE_API_KEY 后即可使用。答案类型由你的 questions 推断;
包内含 ESM、CommonJS 与 TypeScript 声明。三个构造辅助函数:choice()、score()、noul()。
客户端配置 TypeSafeClientConfig
优先级:显式传入值 > 环境变量 > SDK 默认值。
| 选项 | 类型 | 说明与回退 |
|---|---|---|
apiKey | string | 必需的 API key;回退 TYPESAFE_API_KEY |
baseURL | string | API 根;回退 TYPESAFE_BASE_URL → https://api.typesafe.ai |
defaultModel | string | 回退 TYPESAFE_DEFAULT_MODEL → jev-latest |
timeout | number | 每次尝试的超时(毫秒),无总重试预算。默认 10000 |
retry | Partial<RetryPolicy> | 重试覆盖;省略的字段用 RetryPolicy 默认值 |
logLevel | LogLevel | 回退 TYPESAFE_LOG_LEVEL → warn。info 打请求摘要;debug 加 headers 与 body(已知凭据头会被脱敏,body 不会) |
logger | Logger | 按 logLevel 过滤。默认带前缀的 console |
defaultHeaders | Record<string,string> | 额外请求头;单次调用的 header 优先 |
fetch | Fetch | 自定义 fetch,默认全局 fetch |
dangerouslyAllowBrowser | boolean | 允许浏览器端使用,会把 API key 暴露给页面用户。默认 false |
dangerouslyAllowBrowser 字面意思就是「危险地允许浏览器」。前端应用请让后端代理请求。
主要类型
| 类别 | 成员 |
|---|---|
| 核心接口 | SystemOneRequest、SystemOneRequestPayload、SystemOneResult、Questions、Question、ResultFor |
| 问题 / 响应 | ChoiceQuestion、ChoiceCriteria、ChoiceResponse、ScoreQuestion、ScoreCriteria、ScoreLegend、ScoreOf、ScoreResponse、NoulQuestion、NoulResponse |
| 其它 | Usage、WithResponse、RequestOptions、RetryPolicy、ModelCard、Models、EntryType、JsonValue、Description、Logger |
| 常量 | ENV、LOG_LEVELS、VERSION |
Usage 为只读的 input_tokens 与 output_tokens;WithResponse<T> 提供解析后的数据、HTTP 响应与 request ID。
23 · 错误与重试
| 状态码 | 含义 |
|---|---|
401 Unauthorized | 缺少或无效的 API key。检查 Authorization 头。 |
422 Unprocessable Entity | 请求体校验失败 —— 缺少必填字段或问题格式错误。响应体会指出出问题的字段。 |
429 Too Many Requests | 超过速率上限。退避后重试。 |
529 Overloaded | TypeSafe 临时过载。稍后重试。 |
收到 429 或 529 时,用指数退避重试,而不是立刻重试。
官方 SDK 带默认重试策略会自动处理,并在响应携带 retry-after 头时遵守它。
Python 异常层级
TypeSafeError # SDK 失败的基类
├── TypeSafeAPIError # 不成功的 HTTP 响应,带 status / body / headers
├── TypeSafeAPIConnectionError # 请求无法到达或无法读取服务端
└── TypeSafeAPITimeoutError # 请求超过设定的超时
JavaScript 异常层级
TypeSafeError
├── APIError # 基类;带 HTTP status 与 body
│ ├── BadRequestError
│ ├── AuthenticationError
│ ├── PermissionDeniedError
│ ├── NotFoundError
│ ├── RateLimitError
│ ├── UnprocessableEntityError
│ └── InternalServerError
├── APIConnectionError # 请求/响应投递失败(DNS、TLS、连接关闭等)
│ └── APITimeoutError
└── APIUserAbortError # 请求被用户主动中止
APIConnectionError 被 APITimeoutError 继承。
另外 APIPromise<T> 是 SDK 返回的 Promise 类型,
可通过 WithResponse<T> 拿到解析后的数据、HTTP 响应与 request ID。
24 · Cookbook 索引
每份 cookbook 都是一个可跑的完整例子:真实数据集 + 决定某件事的 TypeSafe 问题 + 把这些决定变成可用系统的代码。 阅读前请假定你已掌握三个原语与置信度。
自检一致性(Beginner)
| Cookbook | 做什么 |
|---|---|
| Self-consistency: nouls | 把不确定的概率路由给人工复核,同时保留底层 noul 值可见。 |
| Self-consistency: choices | 给审核决策加一个「不确定」结果,并比较标签一致率与自动执行占比。 |
批处理(Beginner)
| Cookbook | 做什么 |
|---|---|
| Parallel questions | 对 GDPR 维基百科条目跑 13 个问题的合规简报。证明把全部问题合并成一次 TypeSafe 调用比拆开便宜 12.2 倍、快 10 倍,且答案不变。 |
How-to
| Cookbook | 做什么 | 难度 |
|---|---|---|
| Re-ranking | 为 40 条 CLERC 法律查询构建 30 段 BM25 短名单,再用「每个查询-候选对一个 TypeSafe 问题」重排:top-1 准确率 5% → 18%,top-10 38% → 62%。 | Beginner |
| Line-by-line search | 给 GitHub 服务条款做语义搜索。一次请求里用 Choice 对 218 个行 id 打分,并用 Noul 判断文档是否包含答案。 | Beginner |
| Structure recovery | 两次请求把丢失格式的纯文本恢复成 Markdown:一次缝合被硬换行切断的行,一次给每个块分类(标题 / 列表 / 代码 / callout)。 | Beginner |
| Function calling | 把自然语言交易请求变成普通类型化函数的调用 —— 把函数名与闭集参数映射成带置信度的 TypeSafe 问题。 | Intermediate |
| Skill suggestion | 从 Nous Research Hermes 目录的 182 个技能里,为一次智能体回合挑至多一个:两次 TypeSafe 请求,先排序再复核前三名。 | Intermediate |
| Knowledge graph entity alignment | 判断两份啤酒目录的 450 对候选是否描述同一产品:一个 Score + 三个伴随 Noul(用于暴露哪些字段不一致)。 | Beginner |
| Classifying RAG passages | 用一次 TypeSafe 请求给每条召回段落打分,再由代码决定哪些进入作答模型。 | Intermediate |
| Double-checking citations | 对照源文档检查引文是否错误或幻觉:一个 Choice 判断这段引用的上下文是否支撑该主张。 | Beginner |
| Guardrails for LLMs | 用一次 TypeSafe 请求筛查进出 LLM 应用的每条消息,对风险概率与严重度设阈值后放行 / 复核 / 拦截 / 转人工。 | Intermediate |
Extraction
| Cookbook | 做什么 | 难度 |
|---|---|---|
| SDE cascade | 两阶段结构化数据抽取级联(mini → verify → reasoning),以极低成本拿到接近大型推理模型的质量。 | Intermediate |
| Date extraction | 向 TypeSafe 要文档里提到的日期组成部分,再在代码里解析与校验,配合置信度复核,处理绝对与相对日期。 | Beginner |
| Pre-parsed value extraction | 先用正则找出候选的邮箱 / 电话 / 金额,再让 TypeSafe 选出被请求的那一段,代码逐字复制并归一化。 | Beginner |
Classification
| Cookbook | 做什么 | 难度 |
|---|---|---|
| Hierarchical classification | 在专利、零售商品、生物医学与源代码等深层分类体系上做 beam search,遍历 Choice 概率。 | Intermediate |
| Autoresearch feature discovery | 跑一个自动研究循环:提出 TypeSafe 问题 → 把自由文本转成数值特征 → 用模型误差改进一个有监督的 CatBoost 回归器。 | Advanced |
| Classification using confidence | 用每个 SEC 年报一次 Choice 分到 75 个行业组,再读答案自带的置信度决定是报该组还是上报其所属大类。 | Beginner |
「Parallel questions」cookbook 给出的数字是 12.2× 便宜 / 10.0× 快; Primitives 页面引用同一实验时写的是 11.5× / 9.6×。量级一致,但具体比值在两个页面上不同 —— 引用请以 cookbook 自身的数字为准。
25 · Agent Skill
针对 Claude Code、Codex 与其它智能体环境的即装即用技能包。它给编码智能体提供 TypeSafe API 的完整上下文: 三种问题类型、架构模式,以及组织评估任务的最佳实践。
安装
# Claude Code
claude plugin marketplace add typesafe-ai/skills
claude plugin install typesafe@typesafe-ai
# 其他智能体:交互式选择;默认项目局部安装,加 -g 全局安装
npx skills add typesafe-ai/skills --skill typesafe-ai
手动安装:复制 skills/typesafe-ai 整个目录 (含其 reference 文件)到你的技能目录。只用一种安装方式,避免重复副本。
更新
claude plugin marketplace update typesafe-ai
claude plugin update typesafe@typesafe-ai
# 重启 Claude Code 或运行 /reload-plugins
# 自动更新:/plugin → Marketplaces → typesafe-ai → Enable auto-update
# skills.sh 安装方式
npx skills update
示例提示词
在提示词里点名技能(「use the TypeSafe skill」)在任何智能体里都有效。Claude Code 插件还可直接用 /typesafe:typesafe-ai。
Using the TypeSafe skill, explore the project and find opportunities for using
intelligent judgement to stand in for complex parsing or other fragile code.
Using the TypeSafe skill, run some experiments using the TypeSafe API key that I've
exported to `TYPESAFE_API_KEY`. Propose changes based on the most promising results.
Using the TypeSafe skill, analyze my code and see if there are any applicable
cookbooks (https://console.typesafe.ai/docs/cookbooks) that show how I could
refactor my code to be less fragile or complex.
四条 vibe coding 原则
- 先把想法和你的智能体聊透,用上面的提示词起步;
- 实现前先审一遍计划,确认它说得通;
- 把常量(问题、阈值)集中放在一处便于 review。智能体不擅长写问题,预期要和它协作改;
- 不要轻信它的断言,鼓励它验证自己的假设。
常见故障
| 症状 | 怎么办 |
|---|---|
| 智能体没用这个技能 | Claude Code 里用 /typesafe:typesafe-ai;其它智能体说「use the TypeSafe skill」。仍不生效就确认安装器针对的是你正在用的那个智能体,然后重启。 |
| 路由表现不符合预期 | 检查问题与阈值。阈值可能过高(假阴性)或过低(假阳性),也可能需要把问题写得更具体。 |
| 到处都在用置信度阈值 | 如果你只关心挑最好的选项,取最高置信度即可,不必设阈值;如果你心里有特定统计算法,应该用概率而不是置信度。 |
| TypeSafe 代码难 review | 最需要人 review 的是「问题」与「阈值常量」。把它们集中放在单个文件里。 |
| 智能体编造了请求 / 响应字段 | 技能版本旧了。按你的安装方式更新后重试。 |
26 · 最佳实践与反模式
典型决策形状 → 该用什么
| 决策形状 | 什么时候用 | 例子 |
|---|---|---|
| 分类 Classification | 应当有一个已知类别胜出 | 意图、主题、部门、风险类型、实体类型 |
| 检测 Detection | 需要「某个属性存在」的概率 | 垃圾内容、欺诈、紧急性、越狱、敏感数据 |
| 评分 Scoring | 答案落在一个有序 rubric 上 | 严重度、相关性、质量、情绪、适合度 |
| 路由 Routing | 一个类别决定下一段代码路径 | 工具调用、升级、模型路由、支持队列 |
| 搜索 Search | 要找出匹配自然语言查询的条目 | 语义搜索、文档发现、候选生成 |
| 检索 Retrieval | 工作流需要最相关的上下文或记录 | RAG 上下文、证据检索、知识查找 |
| 排序 Ranking | 条目要按语义相关性或质量排 | 搜索结果、推荐、候选优先级 |
| 校验 Verification | 必须检查某个产物有无特定失败模式 | 引文是否有支撑、政策违规、工具调用错误 |
| ML 特征提取 | 下游经典 ML 模型需要语义信号 | 购买意向、产品兴趣、竞争压力、流失信号 |
| 结构化数据抽取 | 必须从非结构化输入里还原已知字段 | 候选人属性、订单字段、文档标签 |
五大用例类别
AI 自动化软件
把 AI 与可靠软件交织在一起,可以在后台跑一百万次而无需人类共同驾驶。代码拥有控制流(不是 markdown 文件),TypeSafe 处理语义决策。
实时应用
150ms 级的前沿智能意味着 AI 的决策可以快过人类感知 —— 快到能被编程去打游戏或嵌入 UI。
大数据上的 AI Map Reduce
便宜 100 倍意味着你能处理巨型数据集:在巨量语料里检索、给海量智能体轨迹分类、抽取特征做预测。
通用校验器
校验任何 AI 的 prompt、抽取结果、推理轨迹、工具调用 —— 检测越狱、引文错误、幻觉,成本只是那次 LLM 调用的零头。
Do(做这些)
- 一次请求打包所有问题,包括投机性的 —— 它们是接近免费的。
- 把宽判断拆成原子问题,在代码里合成。这是整份文档最重要的概念。
- 只给必要上下文。想清楚用条例退税政策时,不要把整个 CRM 导出扔进去。
- question 与 criteria 保持一致,去掉自相矛盾的表述。
- 数学、日期算术、去重留在代码里,Jev 只负责它真正擅长的那部分判断。
- 难分辨的选项用对象描述:
what/not_for/examples。 - 给每一项选项列表加兜底(
other或none of the above)。 - 组合 Score 前先归一化(除以最高等级编号)。
- 把阈值常量集中在一处,方便 review 和调整。
- 钉住版本 ID,如果阈值是针对某个版本标定的。
Don't(别做这些)
- ✗ 每个问题发一次请求 —— 这是编码智能体最容易犯的习惯。
- ✗ 让 Jev 生成文本(可以用 Choice 串联,但这会慢且差)。
- ✗ 依赖实现细节:
P(x)与1-P(not x)不一定相等。 - ✗ 把 Noul 的阈值直接搬到 Choice 上。
- ✗ 把 Noul 的中间值当成「程度中等」。
- ✗ 只看
score而不看probabilities与confidence。 - ✗ 把「写长 prompt」的习惯搬过来 —— prompt 越长,Jev 的原子判断越容易被稀释。
- ✗ 只用 Level 编号「0/1/2」当 criteria 描述 —— 模型没有可比对象,必须配文字定义。
落地检查清单
- 问题是不是原子的?能不能被打断成一件事?
- 是不是只发送了需要的内容?
- criteria 是否互相排斥且覆盖完整?有没有兜底选项?
- 阈值是否写在了一处集中定义的地方?
- 低置信度 / 中等 noul 的路径是转人工 / 追问,而不是硬猜吗?
- 权重是否与 requires 符合?改动它们是否需要重新验证?
- 在真实的业务数据上校准过吗?
内容整理自 docs.typesafe.ai 全站(含 llms-full.txt,共约 20,600 行、
覆盖全部 100+ 页面:概念、原语、模式、HTTP API、Python / JavaScript SDK 全量 API 参考、
15 份 cookbook、模型与锯齿说明)。所有数值示例(概率、置信度、评分、延迟、价格、限额)
均为官方文档给出的实测或标称值,并非我们自己跑出来的 Benchmark。
若与官网存在差异,以官网实时版本为准。