465 lines
16 KiB
Markdown
465 lines
16 KiB
Markdown
# Calculet NPU 新模型适配工程 Runbook
|
||
|
||
日期:2026-08-02
|
||
目标:从一个可公开/合法使用的模型权重,生成可复现、可验证、可服务化的 Calculet NPU 版本
|
||
|
||
## 1. 两阶段策略
|
||
|
||
第一阶段选择 1-7B decoder-only dense 模型,batch 1、单芯粒优先、context 4K/8K。目的不是追求业务最大模型,而是闭环 ONNX、量化、compiler、calbin、Runtime、llama.cpp、golden 和恢复。
|
||
|
||
第二阶段再进入 MoE、双芯粒、40K context、batch 4/8/16。否则一次引入 router、expert kernel、TP 通信、巨型 KV 和 dynamic D2D,任何失败都难以归因。
|
||
|
||
单芯粒产物本身是 C3,必须由厂商 compiler 生成;不能在当前双芯粒 calbin 上只传 `chipId=0`。
|
||
|
||
## 2. 项目角色和职责
|
||
|
||
| 角色 | 负责 | 不可转交的签字 |
|
||
| --- | --- | --- |
|
||
| Model owner | 模型许可、任务质量、参考实现 | 参考输出和质量门槛 |
|
||
| Export owner | HF->ONNX、dynamic/static shape | ONNX checker、数值一致性 |
|
||
| Quant owner | 校准集、quant config、误差 | 量化质量报告 |
|
||
| Compiler/vendor | op lowering、kernel、partition、calbin | 可复现编译和兼容矩阵 |
|
||
| Runtime owner | runner、buffer/KV、错误恢复 | golden、soak、故障注入 |
|
||
| Serving owner | tokenizer、scheduler、API、metrics | 端到端/并发/SLA |
|
||
| Release owner | hash、SBOM、签名、回滚 | release manifest |
|
||
|
||
每个 Gate 有唯一 owner;“厂商支持”不是 owner 名称,必须写到个人/团队和版本交付日期。
|
||
|
||
## 3. 工作目录合同
|
||
|
||
```text
|
||
model-port/<MODEL_ID>/
|
||
00-source/
|
||
source-manifest.json
|
||
hf-config/
|
||
weights/ # 可用引用/对象存储,不要求 Git
|
||
01-reference/
|
||
environment.lock
|
||
prompts.jsonl
|
||
tokens.jsonl
|
||
outputs.jsonl
|
||
tensor-golden/
|
||
02-onnx/
|
||
export-config.yaml
|
||
model.onnx
|
||
model.onnx.data
|
||
onnx-report.json
|
||
03-quant/
|
||
calibration-manifest.jsonl
|
||
quant-config.yaml
|
||
quant-report.json
|
||
04-compiler/
|
||
compiler-manifest.json
|
||
compile-config.yaml
|
||
compile.sh
|
||
logs/
|
||
05-calbin/
|
||
package/
|
||
validation.json
|
||
file-sha256.txt
|
||
06-runtime/
|
||
runner-config.yaml
|
||
correctness/
|
||
performance/
|
||
faults/
|
||
07-serving/
|
||
service-manifest.json
|
||
conformance/
|
||
soak/
|
||
08-release/
|
||
release-manifest.json
|
||
decision.md
|
||
```
|
||
|
||
所有 JSON/YAML 中的路径使用相对路径或内容寻址 URI。绝对编译机路径只保留在受控原始日志,不写进发布 API。
|
||
|
||
## 4. Gate 0:模型立项
|
||
|
||
### 4.1 输入
|
||
|
||
- 模型仓库 URL、revision commit、权重文件 hash。
|
||
- license、商用/再分发约束、依赖许可证。
|
||
- 目标任务、并发、TTFT/TPOT、context、吞吐、功耗和精度要求。
|
||
- 架构 config:层数、hidden、heads/KV heads、head dim、FFN、vocab、RoPE、norm、activation。
|
||
|
||
### 4.2 复杂度打分
|
||
|
||
| 特征 | 0 分 | 1 分 | 3 分 |
|
||
| --- | --- | --- | --- |
|
||
| 架构 | 标准 dense decoder | 小变体 | MoE/multimodal/custom code |
|
||
| attention | MHA/GQA 已有 | sliding window | 新 mask/latent attention |
|
||
| activation/norm | 已支持 | 参数变化 | 自定义算子 |
|
||
| context | <=8K | <=40K | >40K/dynamic |
|
||
| deployment | B1 单芯粒 | 双芯粒 | batch>1/多模型 |
|
||
|
||
首个模型建议总分 <=3。Gate 0 输出 `source-manifest.json` 和 go/no-go;没有许可与产品 SLA 不进入导出。
|
||
|
||
### 4.3 source manifest
|
||
|
||
```json
|
||
{
|
||
"schema":"calculet-model-source/v1",
|
||
"model_id":"vendor/name@revision",
|
||
"license":"SPDX-ID",
|
||
"files":[{"path":"config.json","bytes":0,"sha256":"..."}],
|
||
"architecture":{"layers":32,"hidden":4096,"q_heads":32,"kv_heads":8,"head_dim":128},
|
||
"tokenizer":{"vocab_size":0,"sha256":"..."},
|
||
"targets":{"context":[4096,8192],"batch":[1],"chips":[1]}
|
||
}
|
||
```
|
||
|
||
## 5. Gate 1:CPU/GPU 参考基线
|
||
|
||
### 5.1 环境固定
|
||
|
||
记录 OS、Python、PyTorch/Transformers、CUDA(如有)、模型 revision、tokenizer hash、dtype、seed、determinism flags。保存 lock file 和容器 digest。
|
||
|
||
命令模板(版本仅占位,由项目锁定):
|
||
|
||
```bash
|
||
python -m pip freeze > 01-reference/environment.lock
|
||
sha256sum 00-source/hf-config/* > 01-reference/source-sha256.txt
|
||
python tools/generate_reference.py \
|
||
--model-dir 00-source \
|
||
--prompts 01-reference/prompts.jsonl \
|
||
--dtype float32 \
|
||
--greedy \
|
||
--dump-logits \
|
||
--output 01-reference/outputs.jsonl
|
||
```
|
||
|
||
### 5.2 prompt corpus
|
||
|
||
每条记录含 `sample_id`、原始 text(或受控引用)、token ids、attention mask、position ids、expected generated token ids。覆盖:空/最短、特殊 token、多语种、代码、数字、重复、最大 context 附近、非法超长。
|
||
|
||
### 5.3 输出
|
||
|
||
- FP32 和 BF16 参考 logits/token 序列。
|
||
- 每层或关键边界 tensor golden(输入/输出 hash + stats)。
|
||
- 模型任务质量基线、perplexity/accuracy。
|
||
- tokenizer conformance:encode/decode round-trip、BOS/EOS/chat template。
|
||
|
||
Gate 1 失败条件:同一环境/seed 不能稳定复现;tokenizer 资产不唯一;模型自定义代码来源未审计。
|
||
|
||
## 6. Gate 2:ONNX 导出
|
||
|
||
### 6.1 图边界
|
||
|
||
至少生成两个逻辑 profile:
|
||
|
||
```text
|
||
prefill(input_ids[B,S], position_ids[B,S], empty/past KV)
|
||
-> logits[B,1,V] + KV side effect/outputs
|
||
|
||
decode(input_ids[B,1], position_ids[B,1], past KV up to S-1)
|
||
-> logits[B,1,V] + updated KV
|
||
```
|
||
|
||
当前 Calculet 图将 KV 交给 Runtime manager,最终 ONNX/calcc 如何把 96 个 past K/V 输入转为 KV scatter 是厂商编译契约。不可自行删除 KV 输入而不验证图语义。
|
||
|
||
### 6.2 export config
|
||
|
||
```yaml
|
||
schema: calculet-onnx-export/v1
|
||
opset: VENDOR_CONFIRMED
|
||
external_data: true
|
||
constant_folding: false
|
||
profiles:
|
||
prefill:
|
||
input_ids: [1, 8192]
|
||
position_ids: [1, 8192]
|
||
valid_sequence_lengths: [1, 16, 128, 1024, 8191]
|
||
decode:
|
||
input_ids: [1, 1]
|
||
position_ids: [1, 1]
|
||
past_sequence_max: 8191
|
||
output:
|
||
logits: [1, 1, VOCAB]
|
||
```
|
||
|
||
`opset` 不能凭通用 ONNX 最新版选择,必须用厂商支持矩阵。
|
||
|
||
### 6.3 验收命令模板
|
||
|
||
```bash
|
||
python -m onnx.checker 02-onnx/model.onnx
|
||
python tools/list_onnx_ops.py 02-onnx/model.onnx > 02-onnx/op-inventory.json
|
||
python tools/compare_reference_onnx.py \
|
||
--reference 01-reference \
|
||
--onnx 02-onnx/model.onnx \
|
||
--matrix configs/shape-matrix.yaml \
|
||
--output 02-onnx/onnx-report.json
|
||
```
|
||
|
||
若模型 >2 GiB,external data 文件必须进入 hash manifest。shape inference 后逐 input/output 比对 name/rank/dtype;禁止 exporter 自动改 tokenizer 或 position 逻辑。
|
||
|
||
### 6.4 Gate 2 门槛
|
||
|
||
- ONNX checker success,external data 全部可解析。
|
||
- op inventory 中每个 domain/opset 明确。
|
||
- 100 个固定 prompt 的 greedy tokens 与 Gate 1 一致。
|
||
- 关键 tensor 容差满足 FP32/BF16 基线。
|
||
- prefill/decode 在所有边界 shape 上结果一致。
|
||
|
||
## 7. Gate 3:算子支持差距
|
||
|
||
生成 `op-support.csv`:
|
||
|
||
```text
|
||
node_name,op_type,domain,opset,input_shapes,dtypes,quant,
|
||
support(C0-C4),lowering,kernel,fallback,owner,blocking_issue
|
||
```
|
||
|
||
分类:
|
||
|
||
- Direct:厂商已有 op/kernel。
|
||
- Fuse:必须通过 pass 与相邻 op 融合。
|
||
- Decompose:可降成多个支持 op,但需性能/数值评估。
|
||
- Custom C oplib/kernel:C3。
|
||
- CPU fallback:若存在,必须显式标记并测性能;不允许静默 fallback。
|
||
- Unsupported:阻断编译。
|
||
|
||
每个 unsupported op 给最小复现 ONNX、固定输入、reference output、shape/dtype 列表。厂商不能只回复“支持这个 op”,要返回最低 compiler/runtime 版本和 runnable case。
|
||
|
||
## 8. Gate 4:量化与校准
|
||
|
||
### 8.1 calibration manifest
|
||
|
||
每个样本记录 hash、domain、token count、tokenizer hash、许可/隐私审核。校准集与最终质量测试集必须分离。
|
||
|
||
### 8.2 quant config
|
||
|
||
```yaml
|
||
schema: calculet-quant/v1
|
||
method: VENDOR_TOOL_METHOD
|
||
activation:
|
||
default: int8
|
||
granularity: per_tensor
|
||
calibration: percentile
|
||
weights:
|
||
default: int8
|
||
group_size: null
|
||
overrides:
|
||
lm_head: bfloat16
|
||
router: bfloat16
|
||
sensitive_layers: []
|
||
excluded_ops: []
|
||
seed: 1234
|
||
```
|
||
|
||
当前模型名中的 W8A8/W4AF16 不足以重建量化;必须拿到每 tensor/layer 的规则、scale/zero point 格式、packing、校准算法和工具版本。
|
||
|
||
### 8.3 敏感度流程
|
||
|
||
1. 全 BF16 作为 compiler reference(若硬件/工具支持)。
|
||
2. 全 W8A8 候选。
|
||
3. 分组将 embedding、QKV/O、router、experts、lm_head 回退 BF16,测质量恢复。
|
||
4. 在不敏感 expert/FFN 尝试 W4。
|
||
5. 选择质量、参数大小、带宽、性能 Pareto 点。
|
||
|
||
Gate 4 同时报告 layer error 和任务质量;“校准成功”不等于业务质量达标。
|
||
|
||
## 9. Gate 5:compiler 可复现
|
||
|
||
### 9.1 compiler manifest
|
||
|
||
```json
|
||
{
|
||
"schema":"calculet-compiler/v1",
|
||
"oci_digest":"sha256:...",
|
||
"calcc_commit":"...",
|
||
"tvm_commit":"...",
|
||
"oplib":{"version":"...","sha256":"..."},
|
||
"kernels":{"version":"...","sha256":"..."},
|
||
"license_id":"...",
|
||
"host":{"arch":"x86_64","driver_requirement":"..."}
|
||
}
|
||
```
|
||
|
||
### 9.2 compile config
|
||
|
||
```yaml
|
||
model: ../02-onnx/model.onnx
|
||
onnx_sha256: REQUIRED
|
||
profiles:
|
||
- name: prefill_b1_s8192
|
||
batch: 1
|
||
max_seq: 8192
|
||
is_prefill: true
|
||
- name: decode_b1_s8192
|
||
batch: 1
|
||
max_seq: 8192
|
||
is_prefill: false
|
||
device:
|
||
type: ks1
|
||
chips: 1
|
||
calcores_per_chip: 2
|
||
features:
|
||
flash_attention: true
|
||
one_token: true
|
||
dynamic_d2d: false
|
||
quant_config: ../03-quant/quant-config.yaml
|
||
output: ../05-calbin/package
|
||
```
|
||
|
||
真实 CLI 参数由厂商交付后替换。任何 secret/license server 地址通过环境注入,不写进脚本/日志。
|
||
|
||
### 9.3 可复现门槛
|
||
|
||
- 干净环境加载固定 OCI,单条 `compile.sh` 成功。
|
||
- 同输入至少重编两次。若二进制非 deterministic,要解释时间戳/地址差异,并证明结构/golden 等价。
|
||
- log 含所有 pass、fallback、warning、kernel selection、memory planner 和 partition 决策。
|
||
- 不允许人工复制隐藏参数块或运行未记录命令完成打包。
|
||
|
||
## 10. Gate 6:calbin 静态验收
|
||
|
||
运行:
|
||
|
||
```bash
|
||
ruby calculet_package_validate.rb \
|
||
--full-hash \
|
||
--output 05-calbin/validation.json \
|
||
05-calbin/package
|
||
```
|
||
|
||
检查细节见《calbin 产物字段字典与静态验收规范》。新 dense 单芯粒包预期:`chip_nums=1`、只含 chip0 mask、无跨芯粒 gather/reduce/D2D;但具体文件数量由 compiler schema 决定,不能硬套当前双芯粒 67 文件。
|
||
|
||
Gate 6 必须 ERROR=0;`BLOCKED_REPRODUCIBILITY` 也必须在新模型项目中清零。
|
||
|
||
## 11. Gate 7:独立 Runtime runner
|
||
|
||
### 11.1 runner 接口
|
||
|
||
```text
|
||
calrt_runner inspect --calbin DIR
|
||
calrt_runner golden --calbin DIR --case CASE.json --output RESULT.json
|
||
calrt_runner stress --calbin DIR --matrix MATRIX.yaml --duration 1h
|
||
calrt_runner fault --calbin DIR --scenario wait-timeout
|
||
```
|
||
|
||
runner 不依赖 llama.cpp server,直接覆盖 CreateVDevice/CreateCalbin/configure/buffer/CSR/infer/Wait/KV/release。这样能区分 Runtime/calbin 问题和服务调度问题。
|
||
|
||
### 11.2 最小 golden sequence
|
||
|
||
1. prefill 长度 1/16/17/128/1024/max-1。
|
||
2. 每个 prefill 后 decode 1、32、256 token。
|
||
3. 比较 final logits、greedy token 和 KV length。
|
||
4. free sequence 后资源回收,再用同 seq id。
|
||
5. 两 seq 交替,确认无串扰。
|
||
|
||
### 11.3 故障
|
||
|
||
缺 tensor、错 CSR、越界 slice、KV 满、submit busy、Wait timeout、CCU exception、reset 后 warmup。所有失败必须结构化返回,runner 不 `abort()`。
|
||
|
||
## 12. Gate 8:llama.cpp adapter
|
||
|
||
### 12.1 manifest 驱动
|
||
|
||
adapter 输入一个发布 manifest,不再 substring 找子图:
|
||
|
||
```json
|
||
{
|
||
"generation":1,
|
||
"calbin_sha256":"...",
|
||
"submodels":{
|
||
"prefill_ids":{"exact_name":"...","input0":"inputs[0]","position":"inputs[1]","logits":"outputs[0]"},
|
||
"decode":{"exact_name":"...","input0":"inputs[0]","position":"inputs[1]","logits":"outputs[0]"}
|
||
},
|
||
"csr":{"cur":"cur_seq_len[0]","total":"past_kv_cur_seq_len[0]"},
|
||
"limits":{"batch":1,"context":8192,"vocab":VOCAB}
|
||
}
|
||
```
|
||
|
||
### 12.2 接入检查
|
||
|
||
- tokenizer hash/vocab 与 logits 一致。
|
||
- `n_ctx/n_batch/n_ubatch/n_seq_max` 都来自经过校验的 generation。
|
||
- Runtime tensor layout 不在 adapter 中硬编码 D0。
|
||
- buffer pool 独占 backing storage。
|
||
- 所有 CalRT return code 上抛。
|
||
- 当前不支持的 llama KV copy/shift/state 功能显式禁用。
|
||
|
||
### 12.3 一致性
|
||
|
||
比较独立 runner 与 server 对相同 token ids 的 logits;再比较 HTTP prompt 经 tokenizer 后的 token ids。第一步失败看 adapter/Runtime,第二步失败看 tokenizer/API template。
|
||
|
||
## 13. Gate 9:服务、负载和恢复
|
||
|
||
### 13.1 功能矩阵
|
||
|
||
stream/non-stream、greedy/sampling、stop token/string、logprobs、取消、deadline、超长拒绝、并发 slot、服务 shutdown。未实现的 KV state/prompt sharing 不出现在能力声明中。
|
||
|
||
### 13.2 性能矩阵
|
||
|
||
prompt `1/16/128/512/2K/4K/max-1`,output `1/32/256`,concurrency `1/2/4/8`。batch=1 时并发只是排队/多 slot,不能称硬件 batch。
|
||
|
||
### 13.3 稳定性
|
||
|
||
1h 开发、8h 候选、24h 发布 soak。记录请求成功率、TTFT/TPOT P50/P99、RSS/device memory、KV/pool、温度/功耗、错误/reset。资源在稳定负载后不得单调增长。
|
||
|
||
## 14. Gate 10:发布和回滚
|
||
|
||
`release-manifest.json` 绑定:
|
||
|
||
```text
|
||
source revision + tokenizer hashes
|
||
ONNX hash + export tool versions
|
||
calibration/quant hashes
|
||
compiler OCI/commits/oplib/kernel
|
||
calbin all-file hash manifest
|
||
Runtime/driver/firmware/chip compatibility
|
||
llama.cpp commit/image digest
|
||
correctness/perf/soak/fault report hashes
|
||
known limitations and disabled features
|
||
previous known-good generation
|
||
```
|
||
|
||
发布采用 immutable generation:先静态验证,再 configure/warmup/golden,原子切 alias。回滚不重新编译,直接指向上一已签名 generation;旧 generation 至少保留到新版本完成观察窗口。
|
||
|
||
## 15. Gate 状态表
|
||
|
||
| Gate | 输出 | 硬门槛 | 失败回到 |
|
||
| --- | --- | --- | --- |
|
||
| G0 立项 | source manifest | 许可、目标、复杂度 | 停止/换模型 |
|
||
| G1 参考 | reference corpus | 可复现 | 修环境/tokenizer |
|
||
| G2 ONNX | model+report | checker+数值 | exporter/model wrapper |
|
||
| G3 op gap | support matrix | 无未知 op | vendor/lowering |
|
||
| G4 quant | quant report | 质量达标 | config/calibration |
|
||
| G5 compile | calbin build logs | 可重放 | compiler/vendor |
|
||
| G6 static | validation.json | ERROR=0 | packer/compiler |
|
||
| G7 runner | golden/fault | 正确+可恢复 | Runtime/calbin |
|
||
| G8 adapter | conformance | runner=server | adapter |
|
||
| G9 service | perf/soak | SLA+24h | scheduler/adapter/calbin |
|
||
| G10 release | signed manifest | 可回滚 | 对应失败 Gate |
|
||
|
||
## 16. 第一个项目建议拆票
|
||
|
||
| Ticket | 估算依赖 | 完成定义 |
|
||
| --- | --- | --- |
|
||
| M01 source/tokenizer freeze | model owner | manifest+hash+license |
|
||
| M02 reference harness | model owner | 100 prompt+tensor golden |
|
||
| M03 ONNX exporter | export owner | G2 pass |
|
||
| M04 op support matrix | compiler/vendor | 每 node 有处置 |
|
||
| M05 quant calibration | quant owner | G4 pass |
|
||
| M06 compiler container | vendor | 离线 OCI+SBOM+license |
|
||
| M07 reproducible compile | compiler/vendor | clean build x2 |
|
||
| M08 package validator | Runtime owner | G6 machine pass |
|
||
| M09 CalRT runner | Runtime owner | golden+fault |
|
||
| M10 async/stable adapter | Runtime owner | in-flight=1 24h |
|
||
| M11 service conformance | serving owner | API matrix |
|
||
| M12 performance report | perf owner | raw+summary+CI |
|
||
| M13 release/rollback | release owner | signed generations |
|
||
|
||
## 17. 厂商 kickoff 必须带回的东西
|
||
|
||
- 可加载 compiler OCI,不是远程截图。
|
||
- 真实 compile command/config 和当前 Qwen3 重放样例。
|
||
- ONNX/opset/dynamic shape 支持矩阵。
|
||
- quant/calibration 工具、格式和样例。
|
||
- 单芯粒 dense 示例 calbin、golden、Runtime runner。
|
||
- kernel/fusion/partition/trace 文档和解析器。
|
||
- calbin/Runtime/driver/firmware compatibility matrix。
|
||
- batch、KV、timeout/reset 的精确接口契约。
|
||
|
||
缺前四项时可以继续做 reference/export/harness,但项目状态必须标记 `BLOCKED_COMPILER`,不能承诺新模型上线日期。
|