feat(compute): 补 GLM-5.3 系列规格并标注 Ox Alpha 真身 (#81)

* feat(compute): 补 GLM-5.3 系列规格并标注 Ox Alpha 真身

Ox Alpha 是智谱 GLM-5.3-Flash 的匿名发布别名(canonical_slug
z-ai/glm-5.3-flash-20260826,2026-08-26 揭晓后已从 OpenRouter 列表下架)。

改动:
- model-specs/zhipu.json 新增 glm-5.3 与 glm-5.3-flash 正式规格。两者均
  强制思考(z.ai 文档:thinking.type 仅接受 enabled;OpenRouter models API:
  reasoning.mandatory=true,supported_efforts=[max,high,low]),故
  routing.reasoning.supportedModes 不含 off。按 #79 的教训只用 exact 匹配,
  避免 glm-5.3* 误吞 glm-5.3-flash。
- model-specs/stealth.json 的 ox-alpha 用 spec.extra.modelOrigin 标注真身。
  条目保留:云端算力可能仍以旧别名下发该模型。
- schemas/model-spec.schema.json 为 extra.thinkingOnly / thinkingDefault
  补正式定义与 description。此前这两个键无任何说明,客户端因此从未消费,
  用户选「关闭思考」即触发上游 400(见 desirecore#2307)。
- scripts/validate.mjs 新增静默失效键名巡检:extra 是开放对象,写错键名
  既不报错也无告警。当前巡出 17 处扁平 extra.reasoningEffort。刻意只告警
  不失败——存量取值需逐个核实各自接入面实际接受哪些 effort。

* docs(schema): 写明 extra.reasoning 与 thinkingOnly 的职责边界

provider schema 的 extra.reasoning 此前只有子字段 description、对象本身没有,
维护者看不出它与 model-spec 的 thinkingOnly 分别回答什么问题——issue
desirecore#2307 的误解正源于此。补上对象级说明:本键答「接入面接受哪些 effort
值」且只能写在 provider model;thinkingOnly 答「off 能不能用」、不声明深度档位。
This commit is contained in:
2026-08-27 15:21:25 +08:00
committed by GitHub
parent fa2f10dd0d
commit c46cf07d2f
7 changed files with 196 additions and 7 deletions

View File

@@ -1,6 +1,115 @@
{
"description": "智谱 GLM 系列模型规格。参数来源config-center compute/providers/zhipu.json。glm-5 与 glm-5.1/glm-5-turbo/glm-5v-turbo 前缀相近,故各自仅用 exact 主键匹配,不用宽 pattern 以防误吞。",
"specs": [
{
"id": "glm-5.3",
"displayName": "GLM-5.3",
"family": "glm-5.3",
"match": {
"exact": [
"glm-5.3",
"z-ai/glm-5.3"
]
},
"spec": {
"contextWindow": 1048576,
"maxOutputTokens": 131072,
"capabilities": [
"chat",
"reasoning",
"deep_thinking",
"code",
"math",
"multilingual",
"tool_use",
"agent",
"long_context"
],
"serviceType": [
"chat"
],
"defaultTemperature": 1,
"defaultTopP": 0.95,
"supportsReasoning": true,
"description": "智谱 GLM-5.3 大推理模型,面向复杂软件工程与长程 Agent 任务100 万上下文;强制思考,无法关闭",
"extra": {
"thinkingDefault": true,
"thinkingOnly": true
},
"releasedAt": "2026-08-16"
},
"routing": {
"tier": "flagship",
"routingPriority": 50,
"eligibleForAgent": true,
"defaultReference": false,
"reasoning": {
"supportedModes": [
"auto",
"low",
"high",
"max"
],
"defaultMode": "high"
}
}
},
{
"id": "glm-5.3-flash",
"displayName": "GLM-5.3-Flash",
"family": "glm-5.3-flash",
"match": {
"exact": [
"glm-5.3-flash",
"z-ai/glm-5.3-flash"
]
},
"spec": {
"contextWindow": 1048576,
"maxOutputTokens": 131072,
"capabilities": [
"chat",
"reasoning",
"deep_thinking",
"code",
"tool_use",
"agent",
"long_context",
"vision",
"image_understanding",
"video_understanding"
],
"serviceType": [
"chat",
"reasoning",
"vision"
],
"defaultTemperature": 1,
"defaultTopP": 0.95,
"supportsReasoning": true,
"description": "智谱 GLM-5.3-Flash 原生多模态模型,混合稀疏与线性注意力,面向高效编码与长程 Agent 任务100 万上下文;强制思考,无法关闭",
"extra": {
"thinkingDefault": true,
"thinkingOnly": true
},
"releasedAt": "2026-08-26"
},
"routing": {
"tier": "balanced",
"routingPriority": 48,
"eligibleForAgent": true,
"defaultReference": false,
"reasoning": {
"supportedModes": [
"auto",
"low",
"high",
"max"
],
"defaultMode": "high"
}
}
},
{
"id": "glm-5.2",
"displayName": "GLM-5.2",