- docs/02 架构设计: 新增 aichat审查/异步审批构想/流式渲染调研/generating状态机/密钥迁移健壮性/工作流脚本执行边界/条件表达式引擎/F-07 trait下沉/Agent架构说明/任务推进构想/功能创意池;更新功能决策记录+归档/对抗论证/文档记录规范/经验记录
- docs/03 模块文档: 新增 AI对话引擎/DAG引擎详解;更新 df-knowledge/df-nodes/df-storage/df-workflow/df-ai
- docs/05 代码审查: 新增 全栈审查/全局review/架构审查/近期改动审查/工作区多角度走查/自研memo流式渲染审查
- docs/09 问题排查: 新增 aichat-apikey-401
- docs/INDEX+README 索引同步;docs/todo 待办看板(2026-06-15 汇总)
- PROGRESS.md Sprint 22-25;URGENT.md 加急清单快照(5 项 P0 已全修)
- scripts/cleanup_orphan_tasks.{py,sh} 孤儿任务清理工具
- .gitignore 补 *.broken.bak + tmp/ 噪音排除
142 lines
5.8 KiB
Markdown
142 lines
5.8 KiB
Markdown
# df-storage 存储层
|
||
|
||
> 创建: 2026-06-10 | 最后更新: 2026-06-14
|
||
|
||
---
|
||
|
||
## 概述
|
||
|
||
df-storage 是 DevFlow 的数据持久化层,基于 SQLite (rusqlite),负责连接管理、Schema 迁移(V1-V8)和 CRUD 操作。全部 Repo 由 `impl_repo!` 宏自动生成。
|
||
|
||
---
|
||
|
||
## 当前状态
|
||
|
||
| 功能 | 状态 |
|
||
|------|------|
|
||
| SQLite 连接管理 | ✅ |
|
||
| Schema 迁移 V1-V8 | ✅ |
|
||
| impl_repo! 宏 CRUD | ✅ |
|
||
| KnowledgeRepo(含向量)| ✅ Sprint 15 |
|
||
| 事务支持 | ⬜ 按需 |
|
||
|
||
---
|
||
|
||
## 数据表(V1-V8 迁移历史)
|
||
|
||
| 版本 | 新增/变更 |
|
||
|------|-----------|
|
||
| V1 | ideas / projects / tasks / releases / workflow_executions / node_executions(6 张基础表 + 4 索引)|
|
||
| V2 | ideas 加 promoted_to/ai_analysis/scores;tasks 加 workflow_def_id/base_branch;workflow_executions 加 project_id/task_id;新建 branches 表(含 2 索引)|
|
||
| V3 | ai_providers / ai_conversations / ai_tool_executions(AI 功能 3 张表)|
|
||
| V4 | 幂等补列:ai_conversations.archived(PRAGMA 探测,兼容坏库)|
|
||
| V5 | 幂等补列:ai_conversations.prompt_tokens / completion_tokens / model / models(Token 用量)|
|
||
| V6 | 幂等补列:ai_conversations.skill(技能注入)|
|
||
| V7 | 新建 **knowledges** 表(Sprint 15)|
|
||
| V8 | 幂等补列:knowledges.embedding BLOB(Phase 5.5 向量检索)|
|
||
|
||
### knowledges 表(V7)
|
||
|
||
```sql
|
||
CREATE TABLE IF NOT EXISTS knowledges (
|
||
id TEXT PRIMARY KEY,
|
||
kind TEXT NOT NULL DEFAULT 'pitfall',
|
||
title TEXT NOT NULL,
|
||
content TEXT NOT NULL DEFAULT '',
|
||
tags TEXT, -- JSON array string
|
||
status TEXT NOT NULL DEFAULT 'candidate',
|
||
confidence TEXT, -- 'high'|'medium'|'low'
|
||
reuse_count INTEGER NOT NULL DEFAULT 0,
|
||
verified INTEGER NOT NULL DEFAULT 0, -- 0/1
|
||
source_project TEXT,
|
||
source_ref TEXT,
|
||
created_at TEXT NOT NULL, -- 毫秒字符串
|
||
updated_at TEXT NOT NULL,
|
||
embedding BLOB -- V8 补列,f32 little-endian
|
||
);
|
||
CREATE INDEX IF NOT EXISTS idx_knowledges_status ON knowledges(status);
|
||
CREATE INDEX IF NOT EXISTS idx_knowledges_kind ON knowledges(kind);
|
||
CREATE INDEX IF NOT EXISTS idx_knowledges_reuse_count ON knowledges(reuse_count DESC);
|
||
```
|
||
|
||
---
|
||
|
||
## impl_repo! 宏
|
||
|
||
自动生成以下方法:`insert` / `get_by_id` / `list_all` / `query` / `update_field` / `update_full` / `delete`
|
||
|
||
列名白名单(ALLOWED_COLUMNS)防 SQL 注入,各 Repo 声明各自允许的列。
|
||
|
||
### 已注册 Repo 列表
|
||
|
||
| Repo | 表 |
|
||
|------|----|
|
||
| IdeaRepo | ideas |
|
||
| ProjectRepo | projects |
|
||
| TaskRepo | tasks |
|
||
| ReleaseRepo | releases |
|
||
| WorkflowRepo | workflow_executions |
|
||
| NodeExecutionRepo | node_executions |
|
||
| BranchRepo | branches |
|
||
| AiProviderRepo | ai_providers |
|
||
| AiConversationRepo | ai_conversations |
|
||
| AiToolExecutionRepo | ai_tool_executions |
|
||
| **KnowledgeRepo** | **knowledges** |
|
||
|
||
---
|
||
|
||
## KnowledgeRepo 自定义方法(Sprint 15)
|
||
|
||
| 方法 | 说明 |
|
||
|------|------|
|
||
| `search(query, kind?, limit)` | LIKE 双分支(有/无 kind 过滤),均含 `WHERE status='published'`,`ORDER BY reuse_count DESC LIMIT ?`,top-N≤3 用于注入 |
|
||
| `list_by_status(status)` | CASE WHEN confidence 语义排序(High→Medium→Low),用于审核收件箱 |
|
||
| `increment_reuse_count(id)` | `UPDATE SET reuse_count = reuse_count + 1, updated_at = ?`(SQL 原子操作,连带刷 updated_at)|
|
||
| `top_used(limit)` | published 按 reuse_count DESC,热门列表 |
|
||
| `set_embedding(id, &[f32])` | UPDATE embedding BLOB(f32 little-endian)|
|
||
| `search_vector(query_vec, limit)` | ✅ FR-D2(2026-06-14 commit 4a95f6a):由 `SELECT *` 改为**显式 14 列** `id, kind, title, content, tags, status, confidence, reuse_count, verified, source_project, source_ref, reasoning, created_at, updated_at`;`embedding` 单独另一条 SQL 取(BLOB 单独读,避免与大文本字段混取)→ 纯 Rust 余弦批量比较,skip 维度不匹配。消除 `SELECT *` 隐式依赖,字段精简(`reasoning` 已在列白名单中) |
|
||
| `list_non_archived()` | `WHERE status != 'archived'` 全量,CASE confidence 语义排序(high>medium>low),次 created_at DESC → `Vec<KnowledgeRecord>` |
|
||
|
||
### 向量工具函数
|
||
|
||
```rust
|
||
fn f32s_to_blob(v: &[f32]) -> Vec<u8> // f32 → little-endian bytes
|
||
fn blob_to_f32s(b: &[u8]) -> Vec<f32> // bytes → f32(chunks_exact(4),尾部非 4 倍数残字节截断丢弃)
|
||
fn cosine_similarity(a: &[f32], b: &[f32]) -> f32 // 点积 / (‖a‖‖b‖ + 1e-8 防零除),零向量返回有限值
|
||
```
|
||
|
||
> **不引入 sqlite-vec**:规避 Windows MSVC 下编译 C 扩展的风险。纯 Rust 实现,零外部 C 依赖。
|
||
|
||
---
|
||
|
||
## 迁移幂等设计
|
||
|
||
迁移机制:`schema_version` 表记录当前版本,`MIGRATION_VERSION` 常量为目标版本,`run()` 按 `current_version < N` 顺序应用各版本。
|
||
|
||
### v4 解法:PRAGMA 探测列存在性
|
||
|
||
关键列补建不依赖版本号,用 `PRAGMA table_info(<table>)` 探测实际 schema,缺列才 `ALTER TABLE ADD COLUMN`。
|
||
|
||
- 对新库(列已由建表带入)、老库(DDL 正常生效)、坏库(版本号已写入但 DDL 漏生效)三种情况都安全幂等。
|
||
- V5/V6/V8 均复用此模式。
|
||
|
||
---
|
||
|
||
## 文件结构
|
||
|
||
```
|
||
crates/df-storage/src/
|
||
├── lib.rs — 模块入口,导出公共 API
|
||
├── db.rs — SQLite 连接管理(Database struct)
|
||
├── migrations.rs — V1-V8 迁移逻辑,MIGRATION_VERSION=8
|
||
├── models.rs — 全部 *Record struct(含 KnowledgeRecord)
|
||
└── crud.rs — impl_repo! 宏 + 全部 Repo(含 KnowledgeRepo)
|
||
```
|
||
|
||
---
|
||
|
||
## 相关文档
|
||
|
||
- [SQLite CRUD 模式](../01-技术文档/SQLite-CRUD模式-2026-06-12.md)
|
||
- [前后端类型对齐](../02-架构设计/前后端类型对齐-2026-06-12.md)
|