Files
calculet-npu-research-archive/reports/Calculet-NPU-新模型适配工程Runbook-20260802.md

465 lines
16 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Calculet NPU 新模型适配工程 Runbook
日期:2026-08-02
目标:从一个可公开/合法使用的模型权重,生成可复现、可验证、可服务化的 Calculet NPU 版本
## 1. 两阶段策略
第一阶段选择 1-7B decoder-only dense 模型,batch 1、单芯粒优先、context 4K/8K。目的不是追求业务最大模型,而是闭环 ONNX、量化、compiler、calbin、Runtime、llama.cpp、golden 和恢复。
第二阶段再进入 MoE、双芯粒、40K context、batch 4/8/16。否则一次引入 router、expert kernel、TP 通信、巨型 KV 和 dynamic D2D,任何失败都难以归因。
单芯粒产物本身是 C3,必须由厂商 compiler 生成;不能在当前双芯粒 calbin 上只传 `chipId=0`
## 2. 项目角色和职责
| 角色 | 负责 | 不可转交的签字 |
| --- | --- | --- |
| Model owner | 模型许可、任务质量、参考实现 | 参考输出和质量门槛 |
| Export owner | HF->ONNX、dynamic/static shape | ONNX checker、数值一致性 |
| Quant owner | 校准集、quant config、误差 | 量化质量报告 |
| Compiler/vendor | op lowering、kernel、partition、calbin | 可复现编译和兼容矩阵 |
| Runtime owner | runner、buffer/KV、错误恢复 | golden、soak、故障注入 |
| Serving owner | tokenizer、scheduler、API、metrics | 端到端/并发/SLA |
| Release owner | hash、SBOM、签名、回滚 | release manifest |
每个 Gate 有唯一 owner;“厂商支持”不是 owner 名称,必须写到个人/团队和版本交付日期。
## 3. 工作目录合同
```text
model-port/<MODEL_ID>/
00-source/
source-manifest.json
hf-config/
weights/ # 可用引用/对象存储,不要求 Git
01-reference/
environment.lock
prompts.jsonl
tokens.jsonl
outputs.jsonl
tensor-golden/
02-onnx/
export-config.yaml
model.onnx
model.onnx.data
onnx-report.json
03-quant/
calibration-manifest.jsonl
quant-config.yaml
quant-report.json
04-compiler/
compiler-manifest.json
compile-config.yaml
compile.sh
logs/
05-calbin/
package/
validation.json
file-sha256.txt
06-runtime/
runner-config.yaml
correctness/
performance/
faults/
07-serving/
service-manifest.json
conformance/
soak/
08-release/
release-manifest.json
decision.md
```
所有 JSON/YAML 中的路径使用相对路径或内容寻址 URI。绝对编译机路径只保留在受控原始日志,不写进发布 API。
## 4. Gate 0:模型立项
### 4.1 输入
- 模型仓库 URL、revision commit、权重文件 hash。
- license、商用/再分发约束、依赖许可证。
- 目标任务、并发、TTFT/TPOT、context、吞吐、功耗和精度要求。
- 架构 config:层数、hidden、heads/KV heads、head dim、FFN、vocab、RoPE、norm、activation。
### 4.2 复杂度打分
| 特征 | 0 分 | 1 分 | 3 分 |
| --- | --- | --- | --- |
| 架构 | 标准 dense decoder | 小变体 | MoE/multimodal/custom code |
| attention | MHA/GQA 已有 | sliding window | 新 mask/latent attention |
| activation/norm | 已支持 | 参数变化 | 自定义算子 |
| context | <=8K | <=40K | >40K/dynamic |
| deployment | B1 单芯粒 | 双芯粒 | batch>1/多模型 |
首个模型建议总分 <=3。Gate 0 输出 `source-manifest.json` 和 go/no-go;没有许可与产品 SLA 不进入导出。
### 4.3 source manifest
```json
{
"schema":"calculet-model-source/v1",
"model_id":"vendor/name@revision",
"license":"SPDX-ID",
"files":[{"path":"config.json","bytes":0,"sha256":"..."}],
"architecture":{"layers":32,"hidden":4096,"q_heads":32,"kv_heads":8,"head_dim":128},
"tokenizer":{"vocab_size":0,"sha256":"..."},
"targets":{"context":[4096,8192],"batch":[1],"chips":[1]}
}
```
## 5. Gate 1CPU/GPU 参考基线
### 5.1 环境固定
记录 OS、Python、PyTorch/Transformers、CUDA(如有)、模型 revision、tokenizer hash、dtype、seed、determinism flags。保存 lock file 和容器 digest。
命令模板(版本仅占位,由项目锁定):
```bash
python -m pip freeze > 01-reference/environment.lock
sha256sum 00-source/hf-config/* > 01-reference/source-sha256.txt
python tools/generate_reference.py \
--model-dir 00-source \
--prompts 01-reference/prompts.jsonl \
--dtype float32 \
--greedy \
--dump-logits \
--output 01-reference/outputs.jsonl
```
### 5.2 prompt corpus
每条记录含 `sample_id`、原始 text(或受控引用)、token ids、attention mask、position ids、expected generated token ids。覆盖:空/最短、特殊 token、多语种、代码、数字、重复、最大 context 附近、非法超长。
### 5.3 输出
- FP32 和 BF16 参考 logits/token 序列。
- 每层或关键边界 tensor golden(输入/输出 hash + stats)。
- 模型任务质量基线、perplexity/accuracy。
- tokenizer conformanceencode/decode round-trip、BOS/EOS/chat template。
Gate 1 失败条件:同一环境/seed 不能稳定复现;tokenizer 资产不唯一;模型自定义代码来源未审计。
## 6. Gate 2ONNX 导出
### 6.1 图边界
至少生成两个逻辑 profile
```text
prefill(input_ids[B,S], position_ids[B,S], empty/past KV)
-> logits[B,1,V] + KV side effect/outputs
decode(input_ids[B,1], position_ids[B,1], past KV up to S-1)
-> logits[B,1,V] + updated KV
```
当前 Calculet 图将 KV 交给 Runtime manager,最终 ONNX/calcc 如何把 96 个 past K/V 输入转为 KV scatter 是厂商编译契约。不可自行删除 KV 输入而不验证图语义。
### 6.2 export config
```yaml
schema: calculet-onnx-export/v1
opset: VENDOR_CONFIRMED
external_data: true
constant_folding: false
profiles:
prefill:
input_ids: [1, 8192]
position_ids: [1, 8192]
valid_sequence_lengths: [1, 16, 128, 1024, 8191]
decode:
input_ids: [1, 1]
position_ids: [1, 1]
past_sequence_max: 8191
output:
logits: [1, 1, VOCAB]
```
`opset` 不能凭通用 ONNX 最新版选择,必须用厂商支持矩阵。
### 6.3 验收命令模板
```bash
python -m onnx.checker 02-onnx/model.onnx
python tools/list_onnx_ops.py 02-onnx/model.onnx > 02-onnx/op-inventory.json
python tools/compare_reference_onnx.py \
--reference 01-reference \
--onnx 02-onnx/model.onnx \
--matrix configs/shape-matrix.yaml \
--output 02-onnx/onnx-report.json
```
若模型 >2 GiBexternal data 文件必须进入 hash manifest。shape inference 后逐 input/output 比对 name/rank/dtype;禁止 exporter 自动改 tokenizer 或 position 逻辑。
### 6.4 Gate 2 门槛
- ONNX checker successexternal data 全部可解析。
- op inventory 中每个 domain/opset 明确。
- 100 个固定 prompt 的 greedy tokens 与 Gate 1 一致。
- 关键 tensor 容差满足 FP32/BF16 基线。
- prefill/decode 在所有边界 shape 上结果一致。
## 7. Gate 3:算子支持差距
生成 `op-support.csv`
```text
node_name,op_type,domain,opset,input_shapes,dtypes,quant,
support(C0-C4),lowering,kernel,fallback,owner,blocking_issue
```
分类:
- Direct:厂商已有 op/kernel。
- Fuse:必须通过 pass 与相邻 op 融合。
- Decompose:可降成多个支持 op,但需性能/数值评估。
- Custom C oplib/kernelC3。
- CPU fallback:若存在,必须显式标记并测性能;不允许静默 fallback。
- Unsupported:阻断编译。
每个 unsupported op 给最小复现 ONNX、固定输入、reference output、shape/dtype 列表。厂商不能只回复“支持这个 op”,要返回最低 compiler/runtime 版本和 runnable case。
## 8. Gate 4:量化与校准
### 8.1 calibration manifest
每个样本记录 hash、domain、token count、tokenizer hash、许可/隐私审核。校准集与最终质量测试集必须分离。
### 8.2 quant config
```yaml
schema: calculet-quant/v1
method: VENDOR_TOOL_METHOD
activation:
default: int8
granularity: per_tensor
calibration: percentile
weights:
default: int8
group_size: null
overrides:
lm_head: bfloat16
router: bfloat16
sensitive_layers: []
excluded_ops: []
seed: 1234
```
当前模型名中的 W8A8/W4AF16 不足以重建量化;必须拿到每 tensor/layer 的规则、scale/zero point 格式、packing、校准算法和工具版本。
### 8.3 敏感度流程
1. 全 BF16 作为 compiler reference(若硬件/工具支持)。
2. 全 W8A8 候选。
3. 分组将 embedding、QKV/O、router、experts、lm_head 回退 BF16,测质量恢复。
4. 在不敏感 expert/FFN 尝试 W4。
5. 选择质量、参数大小、带宽、性能 Pareto 点。
Gate 4 同时报告 layer error 和任务质量;“校准成功”不等于业务质量达标。
## 9. Gate 5compiler 可复现
### 9.1 compiler manifest
```json
{
"schema":"calculet-compiler/v1",
"oci_digest":"sha256:...",
"calcc_commit":"...",
"tvm_commit":"...",
"oplib":{"version":"...","sha256":"..."},
"kernels":{"version":"...","sha256":"..."},
"license_id":"...",
"host":{"arch":"x86_64","driver_requirement":"..."}
}
```
### 9.2 compile config
```yaml
model: ../02-onnx/model.onnx
onnx_sha256: REQUIRED
profiles:
- name: prefill_b1_s8192
batch: 1
max_seq: 8192
is_prefill: true
- name: decode_b1_s8192
batch: 1
max_seq: 8192
is_prefill: false
device:
type: ks1
chips: 1
calcores_per_chip: 2
features:
flash_attention: true
one_token: true
dynamic_d2d: false
quant_config: ../03-quant/quant-config.yaml
output: ../05-calbin/package
```
真实 CLI 参数由厂商交付后替换。任何 secret/license server 地址通过环境注入,不写进脚本/日志。
### 9.3 可复现门槛
- 干净环境加载固定 OCI,单条 `compile.sh` 成功。
- 同输入至少重编两次。若二进制非 deterministic,要解释时间戳/地址差异,并证明结构/golden 等价。
- log 含所有 pass、fallback、warning、kernel selection、memory planner 和 partition 决策。
- 不允许人工复制隐藏参数块或运行未记录命令完成打包。
## 10. Gate 6calbin 静态验收
运行:
```bash
ruby calculet_package_validate.rb \
--full-hash \
--output 05-calbin/validation.json \
05-calbin/package
```
检查细节见《calbin 产物字段字典与静态验收规范》。新 dense 单芯粒包预期:`chip_nums=1`、只含 chip0 mask、无跨芯粒 gather/reduce/D2D;但具体文件数量由 compiler schema 决定,不能硬套当前双芯粒 67 文件。
Gate 6 必须 ERROR=0`BLOCKED_REPRODUCIBILITY` 也必须在新模型项目中清零。
## 11. Gate 7:独立 Runtime runner
### 11.1 runner 接口
```text
calrt_runner inspect --calbin DIR
calrt_runner golden --calbin DIR --case CASE.json --output RESULT.json
calrt_runner stress --calbin DIR --matrix MATRIX.yaml --duration 1h
calrt_runner fault --calbin DIR --scenario wait-timeout
```
runner 不依赖 llama.cpp server,直接覆盖 CreateVDevice/CreateCalbin/configure/buffer/CSR/infer/Wait/KV/release。这样能区分 Runtime/calbin 问题和服务调度问题。
### 11.2 最小 golden sequence
1. prefill 长度 1/16/17/128/1024/max-1。
2. 每个 prefill 后 decode 1、32、256 token。
3. 比较 final logits、greedy token 和 KV length。
4. free sequence 后资源回收,再用同 seq id。
5. 两 seq 交替,确认无串扰。
### 11.3 故障
缺 tensor、错 CSR、越界 slice、KV 满、submit busy、Wait timeout、CCU exception、reset 后 warmup。所有失败必须结构化返回,runner 不 `abort()`
## 12. Gate 8llama.cpp adapter
### 12.1 manifest 驱动
adapter 输入一个发布 manifest,不再 substring 找子图:
```json
{
"generation":1,
"calbin_sha256":"...",
"submodels":{
"prefill_ids":{"exact_name":"...","input0":"inputs[0]","position":"inputs[1]","logits":"outputs[0]"},
"decode":{"exact_name":"...","input0":"inputs[0]","position":"inputs[1]","logits":"outputs[0]"}
},
"csr":{"cur":"cur_seq_len[0]","total":"past_kv_cur_seq_len[0]"},
"limits":{"batch":1,"context":8192,"vocab":VOCAB}
}
```
### 12.2 接入检查
- tokenizer hash/vocab 与 logits 一致。
- `n_ctx/n_batch/n_ubatch/n_seq_max` 都来自经过校验的 generation。
- Runtime tensor layout 不在 adapter 中硬编码 D0。
- buffer pool 独占 backing storage。
- 所有 CalRT return code 上抛。
- 当前不支持的 llama KV copy/shift/state 功能显式禁用。
### 12.3 一致性
比较独立 runner 与 server 对相同 token ids 的 logits;再比较 HTTP prompt 经 tokenizer 后的 token ids。第一步失败看 adapter/Runtime,第二步失败看 tokenizer/API template。
## 13. Gate 9:服务、负载和恢复
### 13.1 功能矩阵
stream/non-stream、greedy/sampling、stop token/string、logprobs、取消、deadline、超长拒绝、并发 slot、服务 shutdown。未实现的 KV state/prompt sharing 不出现在能力声明中。
### 13.2 性能矩阵
prompt `1/16/128/512/2K/4K/max-1`output `1/32/256`concurrency `1/2/4/8`。batch=1 时并发只是排队/多 slot,不能称硬件 batch。
### 13.3 稳定性
1h 开发、8h 候选、24h 发布 soak。记录请求成功率、TTFT/TPOT P50/P99、RSS/device memory、KV/pool、温度/功耗、错误/reset。资源在稳定负载后不得单调增长。
## 14. Gate 10:发布和回滚
`release-manifest.json` 绑定:
```text
source revision + tokenizer hashes
ONNX hash + export tool versions
calibration/quant hashes
compiler OCI/commits/oplib/kernel
calbin all-file hash manifest
Runtime/driver/firmware/chip compatibility
llama.cpp commit/image digest
correctness/perf/soak/fault report hashes
known limitations and disabled features
previous known-good generation
```
发布采用 immutable generation:先静态验证,再 configure/warmup/golden,原子切 alias。回滚不重新编译,直接指向上一已签名 generation;旧 generation 至少保留到新版本完成观察窗口。
## 15. Gate 状态表
| Gate | 输出 | 硬门槛 | 失败回到 |
| --- | --- | --- | --- |
| G0 立项 | source manifest | 许可、目标、复杂度 | 停止/换模型 |
| G1 参考 | reference corpus | 可复现 | 修环境/tokenizer |
| G2 ONNX | model+report | checker+数值 | exporter/model wrapper |
| G3 op gap | support matrix | 无未知 op | vendor/lowering |
| G4 quant | quant report | 质量达标 | config/calibration |
| G5 compile | calbin build logs | 可重放 | compiler/vendor |
| G6 static | validation.json | ERROR=0 | packer/compiler |
| G7 runner | golden/fault | 正确+可恢复 | Runtime/calbin |
| G8 adapter | conformance | runner=server | adapter |
| G9 service | perf/soak | SLA+24h | scheduler/adapter/calbin |
| G10 release | signed manifest | 可回滚 | 对应失败 Gate |
## 16. 第一个项目建议拆票
| Ticket | 估算依赖 | 完成定义 |
| --- | --- | --- |
| M01 source/tokenizer freeze | model owner | manifest+hash+license |
| M02 reference harness | model owner | 100 prompt+tensor golden |
| M03 ONNX exporter | export owner | G2 pass |
| M04 op support matrix | compiler/vendor | 每 node 有处置 |
| M05 quant calibration | quant owner | G4 pass |
| M06 compiler container | vendor | 离线 OCI+SBOM+license |
| M07 reproducible compile | compiler/vendor | clean build x2 |
| M08 package validator | Runtime owner | G6 machine pass |
| M09 CalRT runner | Runtime owner | golden+fault |
| M10 async/stable adapter | Runtime owner | in-flight=1 24h |
| M11 service conformance | serving owner | API matrix |
| M12 performance report | perf owner | raw+summary+CI |
| M13 release/rollback | release owner | signed generations |
## 17. 厂商 kickoff 必须带回的东西
- 可加载 compiler OCI,不是远程截图。
- 真实 compile command/config 和当前 Qwen3 重放样例。
- ONNX/opset/dynamic shape 支持矩阵。
- quant/calibration 工具、格式和样例。
- 单芯粒 dense 示例 calbin、golden、Runtime runner。
- kernel/fusion/partition/trace 文档和解析器。
- calbin/Runtime/driver/firmware compatibility matrix。
- batch、KV、timeout/reset 的精确接口契约。
缺前四项时可以继续做 reference/export/harness,但项目状态必须标记 `BLOCKED_COMPILER`,不能承诺新模型上线日期。