- todo: ARC-05越层核验0命中+四领域拆分/ARC-06破环aiShared/FR-D3三子项闭环销账+新增ARC-06b遗留
- 待审查: 登记CR-260616-02(batch29+R-PD-10代码 fc767e1)待审
- DOC-14: df-storage/SQLite-CRUD文档表数V13/V9→V15对齐migrations.rs schema_version=15
8.3 KiB
df-storage 存储层
创建: 2026-06-10 | 最后更新: 2026-06-15
概述
df-storage 是 DevFlow 的数据持久化层,基于 SQLite (rusqlite),负责连接管理、Schema 迁移(V1-V15)和 CRUD 操作。绝大多数 Repo 由 impl_repo! 宏自动生成,KV 类表(app_settings)手写 Repo。
当前状态
| 功能 | 状态 |
|---|---|
| SQLite 连接管理 | ✅ |
| Schema 迁移 V1-V15 | ✅ |
| impl_repo! 宏 CRUD(17 Repo) | ✅ |
| KV 类手写 Repo(SettingsRepo) | ✅ |
| KnowledgeRepo(含向量) | ✅ Sprint 15 |
| KnowledgeEventsRepo(生命线审计) | ✅ |
| 事务支持 | ⬜ 按需 |
数据表(V1-V15 迁移历史)
| 版本 | 新增/变更 |
|---|---|
| V1 | ideas / projects / tasks / releases / workflow_executions / node_executions(6 张基础表 + 4 索引) |
| V2 | ideas 加 promoted_to/ai_analysis/scores;tasks 加 workflow_def_id/base_branch;workflow_executions 加 project_id/task_id;新建 branches 表(含 2 索引) |
| V3 | 新建 ai_conversations 表(archived 列交 v4 幂等补建) |
| V4 | 幂等补列:ai_conversations.archived(PRAGMA 探测,兼容坏库——历史 v3 写入版本号但 ALTER 未生效,统一交 v4 修复) |
| V5 | 幂等补列:ai_conversations.prompt_tokens / completion_tokens(Token 用量) |
| V6 | 幂等补列:ai_conversations.models(对话级多 model 记录) |
| V7 | 新建 knowledges 表(Sprint 15) |
| V8 | 幂等补列:knowledges.embedding BLOB(Phase 5.5 向量检索) |
| V9 | 新建 ai_providers + ai_tool_executions 表(历史遗漏补建,CREATE TABLE IF NOT EXISTS 幂等) |
| V10 | 幂等补列:knowledges.reasoning;新建 knowledge_events 事件表(追加型审计,支撑生命线视图,含 3 索引) |
| V11 | 幂等补列:projects.deleted_at(软删回收站) |
| V12 | 幂等补列:projects.path / projects.stack(项目绑定真实代码目录 + 技术栈) |
| V13 | 新建 app_settings 表(通用应用设置 KV,前端 localStorage 迁移目标) |
| V14 | 幂等补列:tasks.deleted_at(软删回收站,对标 projects.deleted_at V11,D-260616-02 决策) |
| V15 | 幂等补列:tasks.review_rounds INTEGER NOT NULL DEFAULT 0(review 退回累计轮数,F-260616-04 任务推进链) |
当前共 13 张业务表:ideas / projects / tasks / releases / workflow_executions / node_executions / branches / ai_conversations / ai_providers / ai_tool_executions / knowledges / knowledge_events / app_settings。另有一张内部表
schema_version(仅记录版本号,无业务语义)。
knowledges 表(V7 建表 + V8/V10 补列)
CREATE TABLE IF NOT EXISTS knowledges (
id TEXT PRIMARY KEY,
kind TEXT NOT NULL DEFAULT 'pitfall',
title TEXT NOT NULL,
content TEXT NOT NULL DEFAULT '',
tags TEXT, -- JSON array string
status TEXT NOT NULL DEFAULT 'candidate',
confidence TEXT, -- 'high'|'medium'|'low'
reuse_count INTEGER NOT NULL DEFAULT 0,
verified INTEGER NOT NULL DEFAULT 0, -- 0/1
source_project TEXT,
source_ref TEXT,
created_at TEXT NOT NULL, -- 毫秒字符串
updated_at TEXT NOT NULL,
reasoning TEXT, -- V10 补列,AI 提炼「为何值得沉淀」依据
embedding BLOB -- V8 补列,f32 little-endian
);
CREATE INDEX IF NOT EXISTS idx_knowledges_status ON knowledges(status);
CREATE INDEX IF NOT EXISTS idx_knowledges_kind ON knowledges(kind);
CREATE INDEX IF NOT EXISTS idx_knowledges_reuse_count ON knowledges(reuse_count DESC);
impl_repo! 宏
自动生成以下方法:insert / get_by_id / list_all / query / update_field / update_full / delete
列名白名单(ALLOWED_COLUMNS)防 SQL 注入,各 Repo 声明各自允许的列。
KV 类表(
app_settings无固定 schema,按 key 读写)不走impl_repo!宏,由手写SettingsRepo直接 SQL 操作。
已注册 Repo 列表(12 宏生成 + 1 手写 = 13)
| Repo | 表 | 生成方式 |
|---|---|---|
| IdeaRepo | ideas | impl_repo! |
| ProjectRepo | projects | impl_repo! |
| TaskRepo | tasks | impl_repo! |
| ReleaseRepo | releases | impl_repo! |
| WorkflowRepo | workflow_executions | impl_repo! |
| NodeExecutionRepo | node_executions | impl_repo! |
| BranchRepo | branches | impl_repo! |
| AiProviderRepo | ai_providers | impl_repo! |
| AiConversationRepo | ai_conversations | impl_repo! |
| AiToolExecutionRepo | ai_tool_executions | impl_repo! |
| KnowledgeRepo | knowledges | impl_repo!(+ 向量/检索自定义方法) |
| KnowledgeEventsRepo | knowledge_events | impl_repo!(+ list_by_knowledge 等自定义方法) |
| SettingsRepo | app_settings | 手写(KV 表不走宏) |
KnowledgeRepo 自定义方法(Sprint 15)
| 方法 | 说明 |
|---|---|
search(query, kind?, limit) |
LIKE 双分支(有/无 kind 过滤),均含 WHERE status='published',ORDER BY reuse_count DESC LIMIT ?,top-N≤3 用于注入 |
list_by_status(status) |
CASE WHEN confidence 语义排序(High→Medium→Low),用于审核收件箱 |
increment_reuse_count(id) |
UPDATE SET reuse_count = reuse_count + 1, updated_at = ?(SQL 原子操作,连带刷 updated_at) |
top_used(limit) |
published 按 reuse_count DESC,热门列表 |
set_embedding(id, &[f32]) |
UPDATE embedding BLOB(f32 little-endian) |
search_vector(query_vec, limit) |
✅ FR-D2(2026-06-14 commit 4a95f6a):由 SELECT * 改为显式 14 列 id, kind, title, content, tags, status, confidence, reuse_count, verified, source_project, source_ref, reasoning, created_at, updated_at;embedding 单独另一条 SQL 取(BLOB 单独读,避免与大文本字段混取)→ 纯 Rust 余弦批量比较,skip 维度不匹配。消除 SELECT * 隐式依赖,字段精简(reasoning 已在列白名单中) |
list_non_archived() |
WHERE status != 'archived' 全量,CASE confidence 语义排序(high>medium>low),次 created_at DESC → Vec<KnowledgeRecord> |
向量工具函数
fn f32s_to_blob(v: &[f32]) -> Vec<u8> // f32 → little-endian bytes
fn blob_to_f32s(b: &[u8]) -> Vec<f32> // bytes → f32(chunks_exact(4),尾部非 4 倍数残字节截断丢弃)
fn cosine_similarity(a: &[f32], b: &[f32]) -> f32 // 点积 / (‖a‖‖b‖ + 1e-8 防零除),零向量返回有限值
不引入 sqlite-vec:规避 Windows MSVC 下编译 C 扩展的风险。纯 Rust 实现,零外部 C 依赖。
迁移幂等设计
迁移机制:schema_version 表记录已应用版本,run() 读出 MAX(version) 作 current_version,按 current_version < N 顺序应用 V1…V15(每版应用后 INSERT 对应版本号)。当前最新目标版本为 15。
注:早期版本曾用
MIGRATION_VERSION常量,现已无此常量——版本号直接散落在各if current_version < N分支,由schema_version表动态跟踪。
v4 解法:PRAGMA 探测列存在性
关键列补建不依赖版本号,用 PRAGMA table_info(<table>) 探测实际 schema,缺列才 ALTER TABLE ADD COLUMN。
- 对新库(列已由建表带入)、老库(DDL 正常生效)、坏库(版本号已写入但 DDL 漏生效)三种情况都安全幂等。
- V4/V5/V6/V8/V10/V11/V12/V14/V15 均复用此模式(补列迁移)。
- V9/V10/V13 用
CREATE TABLE IF NOT EXISTS幂等建表(老库已有表则跳过)。
文件结构
crates/df-storage/src/
├── lib.rs — 模块入口,导出公共 API
├── db.rs — SQLite 连接管理(Database struct)
├── migrations.rs — V1-V15 迁移逻辑(schema_version 表跟踪版本)
├── models.rs — 全部 *Record struct(含 KnowledgeRecord / KnowledgeEventRecord)
└── crud.rs — impl_repo! 宏 + 全部 Repo(含 KnowledgeRepo / KnowledgeEventsRepo / SettingsRepo)