fix(software): normalize job output artifacts

This commit is contained in:
caoqianming 2026-08-14 08:57:36 +08:00
parent 2560eb4ece
commit 51a213e3a7
14 changed files with 542 additions and 64 deletions

View File

@ -10,6 +10,7 @@
- 新增专业软件 Job 中心Agent 可将当前对话目录内的数据登记为稳定输入,提交、查询和停止 Windows Node 上的受控软件任务;文件栏底部固定展示任务状态,点击可从右侧打开按开启时间倒序、滚动加载的跨对话任务列表,并可收到完成通知或回到原对话分析结果。
- 专业软件任务完成后,输出文件会立即显示在当前对话的文件面板;任务中心收起时也不会再遮挡发送按钮。
- 专业软件生成的正式文件会保存在当前对话的正确目录并记录来源;运行参数和溯源信息收纳到隐藏元数据目录,减少产物列表噪声。
## 0.65.2 — 2026-08-13

View File

@ -56,7 +56,7 @@ zcbot/
**工作目录** = `workspace/users/<user_id>/<working_dir>/`,所有 skill 产物写这里,绝对路径注入 system prompt。`user_id` 走 JWT `sub`,**无 SENTINEL fallback**。DB 内 `name`(显示名)与 `working_dir` 都有值;创建 API 显式给 name 时 working_dir 可留空并沿用 name省略/留空 name 时必须显式给 working_dir、name 先以「新对话」占位并自动生成。二者落库前都是简单名(拒 `/\..`、`.` 起头);同 working_dir 多 task 共享同目录(§7.1)。SaaS 化只换外层根目录,布局不变。
**输出生命周期**(2026-08-06):真实文件仍是事实源,不引入统一产物框架,也不迁移各工具已经稳定的目录。仅对已出现的两类噪声建立局部规则:可过期的外部大响应进用户根目录隐藏缓存 `.zcbot_cache/<task_id>/external_results/`,用户明确要求留存时再 export 到 `<working_dir>/data/external/`;图片/视频的 prompt/model/cost 等技术 sidecar 进产物目录下 `.meta/`。文件面板默认隐藏所有 dotdir用户可在当前 task 工作目录子树内用开关查看 `.meta/`、`.build/`、`.preview/` 等 dotdir但 user_root 的 `.memory/`、`.skills/`、`.kb/`、`.zcbot_*` 永不通过该开关展示,`.env` 等点文件也始终隐藏。`.meta/` 计入用户配额,用户显式查看后可检查、下载或删除,但不自动提升为 artifact。既有可见文件和旧 sidecar 不搬迁、不删除
**输出生命周期**(2026-08-06):真实文件仍是事实源,不引入统一产物框架,也不迁移各工具已经稳定的目录。仅对已出现的局部噪声建立规则:可过期的外部大响应进用户根目录隐藏缓存 `.zcbot_cache/<task_id>/external_results/`,用户明确要求留存时再 export 到 `<working_dir>/data/external/`;图片/视频的 prompt/model/cost,以及专业软件的 plot spec、provenance 等技术 sidecar 进对应产物目录下 `.meta/`。文件面板默认隐藏所有 dotdir用户可在当前 task 工作目录子树内用开关查看 `.meta/`、`.build/`、`.preview/` 等 dotdir但 user_root 的 `.memory/`、`.skills/`、`.kb/`、`.zcbot_*` 永不通过该开关展示,`.env` 等点文件也始终隐藏。`.meta/` 计入用户配额,用户显式查看后可检查、下载或删除,但不自动提升为 artifact。既有普通工具可见文件和旧 sidecar 不做全局搬迁;已知的软件 Job 错位目录通过显式幂等修复脚本校正
| 工作目录位置 | 语义 | Artifact 规则 |
|---|---|---|
@ -173,7 +173,7 @@ Eval 与生产 core 解耦,通过现有 `/v1` API 创建专用任务、监听
**新对话入口(0.60)**:登录未选 task 与左栏「+ 新对话」共用同一前端草稿页,先选择已有 working_dir 或输入新目录名,再直接写消息;草稿不落 DB首发时才 `POST /v1/tasks`,避免空 task 堆积。创建请求省略/留空 name 时必须显式给 working_dir后端据此判定自动命名以「新对话」占位并置一次性 `auto_title_pending`;显式 name 的旧调用继续视为人工标题working_dir 仍可省略并 fallback 到 name`auto_title` 字段只作兼容保留。首条消息并行触发短标题调用,结果只改 `tasks.name`、绝不改 working_dir人工 PATCH name 同时清 pending条件 UPDATE 保证在途标题也不能覆盖用户命名。原完整创建表单保留为「自定义」入口UI 同样要求明确选择 working_dirname 可选,并可预设 description/skill/model。标题是 UI 元数据辅助调用,记 `usage_events.kind="task_title"`;模型调用失败时以首条消息第一行生成本地兜底标题,不阻塞主 run也不把 pending 留给后续消息误命名。
**对话产物与生命周期(0025/0028)**:真实文件仍是内容事实源;`artifacts` 表记录已发布产物的稳定身份和生命周期,包含 user-root 相对当前路径、来源 task、复制来源、哈希/大小及 active/deleted、回收路径。新 `messages.artifact_refs` 使用 `{version:2, artifact_id, scope:"working_dir", path:"reports/a.pdf", label?:"最终报告"}``path` 是兼容快照,预览/下载优先按 `artifact_id` 找当前路径因此移动或重命名后历史卡片仍有效。version 1 和 `NULL` 旧消息继续走原 task-scoped 兼容链。普通源码树、中间文件和配套资源不登记agent 仅用 `publish_artifacts` 显式提升少量最终文件。移动保持身份,复制为每个副本创建新身份并记录直接来源;删除将文件移入 `.zcbot_artifact_trash/` 并软删记录,普通文件仍物理删除。
**对话产物与生命周期(0025/0028/0033)**:真实文件仍是内容事实源;`artifacts` 表记录已发布产物的稳定身份和生命周期,包含 user-root 相对当前路径、来源 task、复制来源、可空的 `software_job_id`哈希/大小及 active/deleted、回收路径。`software_job_id` 不设外键,非空即表示该正式产物由对应专业软件 Job 生成复制品保留该来源Job 或 task 生命周期结束也不抹除来源标识。新 `messages.artifact_refs` 使用 `{version:2, artifact_id, scope:"working_dir", path:"reports/a.pdf", label?:"最终报告"}``path` 是兼容快照,预览/下载优先按 `artifact_id` 找当前路径因此移动或重命名后历史卡片仍有效。version 1 和 `NULL` 旧消息继续走原 task-scoped 兼容链。普通源码树、中间文件、`.meta/` 技术信息和配套资源不登记agent 仅用 `publish_artifacts` 显式提升少量最终文件。移动保持身份,复制为每个副本创建新身份并记录直接来源;删除将文件移入 `.zcbot_artifact_trash/` 并软删记录,普通文件仍物理删除。
**用户消息附件(0031)**:`messages.attachment_refs` 与助手产物分开表达输入附件,元素为 `{version:1, scope:"working_dir", path, label, kind, media_type, size_bytes}`;文件仍是事实源,不登记为已发布 artifact也不承诺独立生命周期。`payload.content` 只保存用户自然语言,模型上下文在内存中按附件类型补兼容提示,避免 UI 协议污染正文。`NULL` 表示旧客户端/旧历史,前端继续解析正文标记;`[]` 表示新消息明确无附件。独立 attachment 表暂缓,只有出现跨消息复用、稳定身份、版本快照或附件级审计需求时再抽象 `file_assets + message_attachments`
@ -470,7 +470,7 @@ Node 通过 `Authorization: Bearer` 与 `X-Node-Id` 建立 `/v1/software-nodes/c
第四阶段落地固定 Origin WorkerNode 仅从管理员安装的固定 Python 运行时启动随程序发布的 `worker.py`,参数只有本机 job 目录请求不能指定脚本、解释器或文件路径。Worker 使用 `originpro` 生成 OPJU、PNG、SVG、PDF、plot spec 和 provenance校验产物签名并原子写入终态当前受控图形仅含 line、scatter、line_scatter 和双栏出版布局。进程内 pipeline 按 job 去重,并脱离单次 WebSocket 的取消令牌运行;连接中断只延迟状态/终态上报。Node 进程若在 Worker 启动后重启,则保守失败而不重复驱动 Origin避免无法证明的双执行。
第五阶段完成输出上传与发布Node 只按固定 manifest ID 逐项流式 PUT并携带 Node、lease、request digest 与内容摘要;云端重新绑定任务身份,不信 Node 提供的路径或媒体类型。文件先进入用户根下隐藏暂存区,固定文件名、单文件/总大小和 SHA-256 全部验证后,目录原子移动到 `<working_dir>/origin/<job_id>/`,再登记平台 artifact UUID 并写成功终态。重复 PUT、complete 和重连均按摘要幂等;部分上传不可见,只有完整集合才能发布。
第五阶段完成输出上传与发布Node 只按固定 manifest ID 逐项流式 PUT并携带 Node、lease、request digest 与内容摘要;云端重新绑定任务身份,不信 Node 提供的路径或媒体类型。文件先进入用户根下隐藏暂存区,固定文件名、单文件/总大小和 SHA-256 全部验证后,把 plot spec、provenance 整理进 `.meta/`,再将完整目录原子移动到 `<working_dir>/origin/<job_id>/`。PNG/SVG/PDF/OPJU 等正式输出登记平台 artifact UUID 和 `software_job_id``.meta/` 只落真实文件;成功状态返回 task-relative `output_dir`Agent 以该目录为起点按需搜索。重复 PUT、complete 和重连均按摘要幂等;部分上传不可见,只有完整集合才能发布。
第六阶段增加用户级 Job 中心与 Agent typed tools。`software_capability_list` 只暴露固定能力及当前在线空闲节点数,`software_job_submit/status/cancel` 在构造时绑定当前 user/task模型不能跨用户或跨对话指定归属。`software_job_submit` 的模型可见唯一入口为 `input_id + plot + output``register_artifact` 是普通文件获得输入身份的唯一入口,并明确返回 UUID。内部完整 `request` 形式仅保留执行层兼容,不进入工具 schema。右下角 Job 中心按用户聚合各对话任务,活动期短轮询、空闲期降频;终态变化通知用户,成功任务可回到原对话发起分析。取消采用协作协议:未派发任务直接终止,已派发任务先进入 `cancelling`,云端通过 WebSocket 发送并在心跳时重放 `job_cancel`Node 杀死固定 Worker 进程树后回报 `cancelled`;终态写入仍由云端账本裁决。

View File

@ -2,7 +2,7 @@
> 配合 `DESIGN.md`。本文件只记 phase 状态、决策偏差、文件量、下一步。每条 1-2 句:做了啥 + 关键判断;细节查 `git log` / `git diff` / `DESIGN §7.9`
最后更新:2026-08-13(专业软件 Job 中心开发中,未发版)
最后更新:2026-08-14(专业软件输出归一与来源标记开发中,未发版)
---
@ -20,6 +20,10 @@
---
## 已完成关键能力
### 2026-08-14
- **08-14 / Unreleased / 专业软件输出目录归一 + artifact 来源**:修正 `tasks.working_dir` 被二次拼入 user root 导致 Job 输出错位的问题,正式输出原子发布到 `<working_dir>/origin/<job_id>/` 并以 nullable `software_job_id` 记录来源plot spec/provenance 改入 `.meta/` 且不再登记 artifact。新增 0033 migration 和只认显式迁移库地址、默认 dry-run、冲突即停的存量目录修复脚本相关 Python 63 项、Node 前端 26 项、编译、Alembic 单 head、Ruff 致命规则及 diff 检查通过,未连接或写入生产 DB。
### 2026-08-13
- **08-13 / Unreleased / 专业软件输入 artifact 契约**:新增 `register_artifact`,将 task 内普通文件登记为稳定输入但不发布聊天交付卡片;`software_job_submit` 改为模型可见的完整 Origin schema并强制使用 artifact UUID已有 artifact 可直接复用,内部完整 request 仅保留兼容。同步校准当前三种图形类型与运行/设计文档;相关 51 项 unittest、Python 编译和 diff 检查通过,未连接或写入生产 DB。

14
RUN.md
View File

@ -1054,7 +1054,15 @@ sudo xfs_quota -x -c "limit -p bhard=10g zcbot_<user_uuid>" /opt
### Windows Node 内网 MVP开发中
先执行 `alembic upgrade head` 创建 `software_node_enrollments`、`software_nodes` 和 `software_jobs`。不要在未确认目标数据库时运行迁移;本机 `.env``ZCBOT_DB_URL` 可能是生产隧道。
先执行 `.venv/Scripts/python.exe main.py db upgrade head` 创建 `software_node_enrollments`、`software_nodes`、`software_jobs`,并为 artifact 增加专业软件来源字段。不要在未确认目标数据库时运行迁移;本机 `.env``ZCBOT_DB_URL` 可能是生产隧道。
0032 旧版曾把 `tasks.working_dir` 二次拼进 user root已完成 Job 的文件和 artifact 路径可能位于重复的 `workspace/users/<uid>/...` 子树。升级到 0033 后,用显式迁移库地址先 dry-run确认所有目标目录无冲突后再加 `--apply`。脚本不加载 `.env`,文件和数据库更新均按 Job manifest 精确处理,目标目录同时存在时会停止,不能覆盖。
```powershell
$env:ZCBOT_MIGRATION_DB_URL="postgresql+psycopg://user:pass@host:5432/zcbot"
.venv/Scripts/python.exe scripts/repair_software_job_outputs.py
.venv/Scripts/python.exe scripts/repair_software_job_outputs.py --apply
```
云端当前提供:
@ -1090,9 +1098,9 @@ cd /d D:\ZcbotNode
install-windows-node.bat
```
若 Python 未加入 PATH可把绝对路径作为第一个参数例如 `install-windows-node.bat "C:\Python312\python.exe"`。默认解释器为 `%ProgramData%\Zcbot\WindowsNode\runtimes\origin\Scripts\python.exe`。如需使用其他受管解释器,设置机器级 `ZCBOT_ORIGIN_PYTHON` 为绝对 `python.exe` 路径后重启 Node。`node.json`、可恢复任务和 runtime 集中保存在 `%ProgramData%\Zcbot\WindowsNode\`,不会因替换程序目录而丢失。运行时固定依赖见发布目录的 `origin-worker/requirements.txt`;任务请求无权选择解释器、脚本或路径。当前 Worker 支持 CSV/XLSX/JSON 输入,`line`、`scatter`、`line_scatter` 与 OPJU/PNG/SVG/PDF 输出。成功产物由 Node 流式上传,全部校验通过后发布到任务工作目录 `origin/<job_id>/`;上传中断会在重连时幂等续传。
若 Python 未加入 PATH可把绝对路径作为第一个参数例如 `install-windows-node.bat "C:\Python312\python.exe"`。默认解释器为 `%ProgramData%\Zcbot\WindowsNode\runtimes\origin\Scripts\python.exe`。如需使用其他受管解释器,设置机器级 `ZCBOT_ORIGIN_PYTHON` 为绝对 `python.exe` 路径后重启 Node。`node.json`、可恢复任务和 runtime 集中保存在 `%ProgramData%\Zcbot\WindowsNode\`,不会因替换程序目录而丢失。运行时固定依赖见发布目录的 `origin-worker/requirements.txt`;任务请求无权选择解释器、脚本或路径。当前 Worker 支持 CSV/XLSX/JSON 输入,`line`、`scatter`、`line_scatter` 与 OPJU/PNG/SVG/PDF 输出。成功产物由 Node 流式上传,全部校验通过后发布到任务工作目录 `origin/<job_id>/`plot spec 与 provenance 位于其 `.meta/`;上传中断会在重连时幂等续传。
Web 用户登录后,右下角 Job 中心会聚合本人最近任务。活动任务约 4 秒刷新一次,空闲时降为约 30 秒停止已派发任务是协作取消状态先显示“正在停止”Node 在线时立即接收断线后在下次连接或心跳时重放。Agent 可调用 `software_capability_list`、`register_artifact`、`software_job_submit`、`software_job_status` 和 `software_job_cancel`。Origin 输入必须是 artifact已有 UUID 可直接提交,普通 task 文件先用相对路径登记;登记不会发布聊天交付卡片。提交工具接收 `input_id`、`plot`、`output`,只创建固定 schema 的持久任务,不会阻塞当前对话等待完成。
Web 用户登录后,文件栏 Job 中心会聚合本人最近任务。活动任务约 4 秒刷新一次,空闲时降为约 30 秒停止已派发任务是协作取消状态先显示“正在停止”Node 在线时立即接收断线后在下次连接或心跳时重放。Agent 可调用 `software_capability_list`、`register_artifact`、`software_job_submit`、`software_job_status` 和 `software_job_cancel`。Origin 输入必须是 artifact已有 UUID 可直接提交,普通 task 文件先用相对路径登记;登记不会发布聊天交付卡片。提交工具接收 `input_id`、`plot`、`output`,只创建固定 schema 的持久任务,不会阻塞当前对话等待完成。成功状态提供 `output_dir`Agent可在该目录内搜索并分析正式输出的 artifact 带 `software_job_id`,供结果卡和产物详情展示来源。
注册配置写入 `%ProgramData%\Zcbot\WindowsNode\node.json`Token 使用 DPAPI `LocalMachine` 加密ACL 仅允许注册账号和 `SYSTEM`。应始终用同一专用 Windows 账号执行统一安装器、注册并运行 Node。当前 MVP 以该账号的登录后计划任务启动,不安装 Windows Service。

View File

@ -46,6 +46,7 @@ def register_published_artifacts(
working_dir: Path,
refs: tuple[dict, ...],
preserve_existing_labels: bool = False,
software_job_id: UUID | None = None,
) -> tuple[dict, ...]:
"""Upsert active artifact identities and return version-2 message refs."""
root = Path(user_root).resolve()
@ -71,9 +72,12 @@ def register_published_artifacts(
}
if not preserve_existing_labels:
update_values["label"] = label
if software_job_id is not None:
update_values["software_job_id"] = software_job_id
statement = pg_insert(Artifact).values(
user_id=user_id,
origin_task_id=task_id,
software_job_id=software_job_id,
current_path=current_path,
label=label,
media_type=media_type,
@ -186,6 +190,7 @@ def copy_active_artifacts(
user_id=user_id,
origin_task_id=original.origin_task_id,
copied_from_artifact_id=original.artifact_id,
software_job_id=original.software_job_id,
current_path=copied_path,
label=(original.label + " 副本").strip(),
media_type=original.media_type,

View File

@ -36,6 +36,16 @@ OUTPUT_ARTIFACTS = {
"plot_spec": ("plot-spec.json", "application/json", None),
"provenance": ("provenance.json", "application/json", None),
}
SOFTWARE_JOB_METADATA_IDS = frozenset({"plot_spec", "provenance"})
def software_job_output_path(output_id: str) -> str:
"""返回 Job 输出目录内路径;技术元数据固定进入隐藏 `.meta/`。"""
item = OUTPUT_ARTIFACTS.get(output_id)
if item is None:
raise SoftwareJobError("unsupported output artifact identity")
filename = item[0]
return f".meta/{filename}" if output_id in SOFTWARE_JOB_METADATA_IDS else filename
def _has_only(value: dict, fields: set[str]) -> bool:
@ -151,6 +161,7 @@ def _job_dict(row: SoftwareJob) -> dict:
"metrics": row.metrics,
"error": row.error,
"artifact_manifest": row.artifact_manifest,
"output_dir": f"origin/{row.job_id}",
"created_at": row.created_at.isoformat() if row.created_at else None,
"started_at": row.started_at.isoformat() if row.started_at else None,
"terminal_at": row.terminal_at.isoformat() if row.terminal_at else None,
@ -695,10 +706,7 @@ def record_job_terminal(node_id: UUID, payload: dict) -> None:
],
)
if len(expected) != len(manifest) or any(
not isinstance(item.get("artifact_id"), str)
or not _is_uuid(item["artifact_id"])
or not isinstance(item.get("path"), str)
or not item["path"].startswith(f"origin/{job.job_id}/")
not _published_output_is_valid(job.job_id, item)
for item in manifest
):
raise SoftwareJobError("successful job artifacts have not been published")
@ -724,6 +732,23 @@ def _is_uuid(value: str) -> bool:
return True
def _published_output_is_valid(job_id: UUID, item: object) -> bool:
if not isinstance(item, dict):
return False
output_id = item.get("source_artifact_id")
if not isinstance(output_id, str):
return False
if output_id not in OUTPUT_ARTIFACTS:
return False
expected_path = f"origin/{job_id}/{software_job_output_path(output_id)}"
if item.get("path") != expected_path:
return False
artifact_id = item.get("artifact_id")
if output_id in SOFTWARE_JOB_METADATA_IDS:
return artifact_id is None
return isinstance(artifact_id, str) and _is_uuid(artifact_id)
def mark_node_jobs_disconnected(node_id: UUID) -> None:
"""连接丢失后保留 Node 归属和 lease禁止任务被自动重派。"""
with session_scope() as session:

View File

@ -222,6 +222,7 @@ class Artifact(Base):
__table_args__ = (
Index("ix_artifacts_user_status_path", "user_id", "status", "current_path"),
Index("ix_artifacts_origin_task", "origin_task_id"),
Index("ix_artifacts_software_job_id", "software_job_id"),
)
artifact_id: Mapped[UUID] = mapped_column(
@ -240,6 +241,10 @@ class Artifact(Base):
ForeignKey("artifacts.artifact_id", ondelete="SET NULL"),
nullable=True,
)
# 不设外键:软件 Job 的生命周期可能短于被复制或长期保留的产物。
software_job_id: Mapped[Optional[UUID]] = mapped_column(
PG_UUID(as_uuid=True), nullable=True
)
current_path: Mapped[str] = mapped_column(Text, nullable=False)
label: Mapped[str] = mapped_column(Text, nullable=False, default="")
media_type: Mapped[Optional[str]] = mapped_column(Text, nullable=True)

View File

@ -0,0 +1,45 @@
"""Record the professional software job that produced an artifact.
Revision ID: 0033
Revises: 0032
Create Date: 2026-08-14
"""
from collections.abc import Sequence
import sqlalchemy as sa
from alembic import op
from sqlalchemy.dialects import postgresql
revision: str = "0033"
down_revision: str | None = "0032"
branch_labels: str | Sequence[str] | None = None
depends_on: str | Sequence[str] | None = None
def upgrade() -> None:
op.add_column(
"artifacts",
sa.Column("software_job_id", postgresql.UUID(as_uuid=True), nullable=True),
)
op.create_index(
"ix_artifacts_software_job_id", "artifacts", ["software_job_id"]
)
# 0032 已发布任务的 manifest 保存了平台 artifact UUID可无损补来源。
op.execute(
sa.text(
"""
UPDATE artifacts AS artifact
SET software_job_id = job.job_id
FROM software_jobs AS job,
LATERAL jsonb_array_elements(job.artifact_manifest) AS output
WHERE output->>'artifact_id' = artifact.artifact_id::text
AND output->>'source_artifact_id' NOT IN ('plot_spec', 'provenance')
AND artifact.software_job_id IS NULL
"""
)
)
def downgrade() -> None:
op.drop_index("ix_artifacts_software_job_id", table_name="artifacts")
op.drop_column("artifacts", "software_job_id")

View File

@ -0,0 +1,197 @@
"""Repair software-job outputs written under the duplicated legacy user path.
The command is dry-run by default. It reads its database URL only from the
explicit ``ZCBOT_MIGRATION_DB_URL`` environment variable and never loads .env.
Run schema migration 0033 before applying this repair.
"""
# ruff: noqa: I001
from __future__ import annotations
import argparse
import os
import sys
from dataclasses import dataclass
from pathlib import Path
from uuid import UUID
from sqlalchemy import create_engine, select
from sqlalchemy.orm import Session
from sqlalchemy.orm.attributes import flag_modified
ROOT = Path(__file__).resolve().parents[1]
sys.path.insert(0, str(ROOT))
from core.software_jobs import (
SOFTWARE_JOB_METADATA_IDS,
software_job_output_path,
)
from core.storage.models import Artifact, SoftwareJob, Task
from core.paths import from_db_path
@dataclass(frozen=True)
class RepairPlan:
job_id: UUID
user_root: Path
working_dir: Path
legacy_dir: Path
output_dir: Path
move_legacy_dir: bool
metadata_moves: tuple[tuple[Path, Path], ...]
def _user_root(working_dir: Path, user_id: UUID) -> Path:
expected = str(user_id)
for candidate in (working_dir, *working_dir.parents):
if candidate.name == expected and candidate.parent.name == "users":
return candidate.resolve()
raise RuntimeError(f"task working_dir has no user root for {user_id}")
def _within(root: Path, target: Path) -> Path:
resolved = target.resolve()
resolved.relative_to(root.resolve())
return resolved
def build_plan(job: SoftwareJob, task: Task) -> RepairPlan | None:
working_dir = from_db_path(task.working_dir).resolve()
user_root = _user_root(working_dir, job.user_id)
output_dir = _within(user_root, working_dir / "origin" / str(job.job_id))
legacy_dir = _within(
user_root,
user_root / Path(task.working_dir) / "origin" / str(job.job_id),
)
legacy_exists = legacy_dir.is_dir()
output_exists = output_dir.is_dir()
if legacy_exists and output_exists:
raise RuntimeError(f"job {job.job_id}: legacy and target output directories both exist")
if not legacy_exists and not output_exists:
print(f"[WARN] job {job.job_id}: output directory is missing")
return None
source_dir = legacy_dir if legacy_exists else output_dir
metadata_moves: list[tuple[Path, Path]] = []
for item in job.artifact_manifest or []:
output_id = str((item or {}).get("source_artifact_id") or "")
if output_id not in SOFTWARE_JOB_METADATA_IDS:
continue
source = source_dir / str((item or {}).get("filename") or "")
destination = source_dir / software_job_output_path(output_id)
if source == destination or not source.exists():
continue
if destination.exists():
raise RuntimeError(f"job {job.job_id}: metadata destination already exists")
metadata_moves.append((source, destination))
return RepairPlan(
job_id=job.job_id,
user_root=user_root,
working_dir=working_dir,
legacy_dir=legacy_dir,
output_dir=output_dir,
move_legacy_dir=legacy_exists,
metadata_moves=tuple(metadata_moves),
)
def apply_files(plan: RepairPlan) -> None:
for _, destination in plan.metadata_moves:
destination.parent.mkdir(parents=True, exist_ok=True)
for source, destination in plan.metadata_moves:
os.replace(source, destination)
if plan.move_legacy_dir:
plan.output_dir.parent.mkdir(parents=True, exist_ok=True)
os.replace(plan.legacy_dir, plan.output_dir)
def update_rows(session: Session, job: SoftwareJob, task: Task, plan: RepairPlan) -> int:
artifact_count = 0
manifest = [dict(item) for item in (job.artifact_manifest or [])]
task_prefix = plan.working_dir.relative_to(plan.user_root).as_posix()
for item in manifest:
output_id = str(item.get("source_artifact_id") or "")
if not output_id:
continue
path = f"origin/{job.job_id}/{software_job_output_path(output_id)}"
item["path"] = path
raw_artifact_id = item.get("artifact_id")
if not raw_artifact_id:
continue
try:
artifact_id = UUID(str(raw_artifact_id))
except ValueError:
continue
artifact = session.get(Artifact, artifact_id)
if artifact is None or artifact.user_id != job.user_id:
raise RuntimeError(f"job {job.job_id}: artifact {artifact_id} is missing")
artifact.current_path = f"{task_prefix}/{path}"
if output_id not in SOFTWARE_JOB_METADATA_IDS:
artifact.software_job_id = job.job_id
artifact_count += 1
job.artifact_manifest = manifest
flag_modified(job, "artifact_manifest")
return artifact_count
def repair(session: Session, *, apply: bool, job_id: UUID | None = None) -> tuple[int, int]:
statement = (
select(SoftwareJob, Task)
.join(Task, Task.task_id == SoftwareJob.task_id)
.where(SoftwareJob.status == "succeeded")
.order_by(SoftwareJob.created_at, SoftwareJob.job_id)
)
if job_id is not None:
statement = statement.where(SoftwareJob.job_id == job_id)
rows = session.execute(statement).all()
prepared: list[tuple[SoftwareJob, Task, RepairPlan]] = []
for job, task in rows:
plan = build_plan(job, task)
if plan is None:
continue
print(
f"[INFO] job={job.job_id} move_dir={plan.move_legacy_dir} "
f"metadata_moves={len(plan.metadata_moves)} target={plan.output_dir}"
)
prepared.append((job, task, plan))
if not apply:
return len(prepared), 0
updated_artifacts = 0
for job, task, plan in prepared:
apply_files(plan)
updated_artifacts += update_rows(session, job, task, plan)
session.flush()
return len(prepared), updated_artifacts
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--apply", action="store_true", help="apply filesystem and database changes")
parser.add_argument("--job-id", type=UUID, help="limit repair to one software job")
args = parser.parse_args()
database_url = os.environ.get("ZCBOT_MIGRATION_DB_URL", "").strip()
if not database_url:
print("[ERR] ZCBOT_MIGRATION_DB_URL is required", file=sys.stderr)
return 2
engine = create_engine(database_url, pool_pre_ping=True, future=True)
try:
with Session(engine, future=True) as session:
try:
jobs, artifacts = repair(
session, apply=args.apply, job_id=args.job_id
)
if args.apply:
session.commit()
else:
session.rollback()
except Exception:
session.rollback()
raise
finally:
engine.dispose()
mode = "applied" if args.apply else "validated"
print(f"[OK] {mode} jobs: {jobs}")
print(f"[OK] updated artifacts: {artifacts}")
return 0
if __name__ == "__main__":
raise SystemExit(main())

View File

@ -0,0 +1,53 @@
from __future__ import annotations
import tempfile
import unittest
from pathlib import Path
from types import SimpleNamespace
from unittest.mock import patch
from uuid import uuid4
from scripts.repair_software_job_outputs import apply_files, build_plan
class RepairSoftwareJobOutputsTests(unittest.TestCase):
def test_legacy_directory_moves_and_metadata_becomes_hidden(self) -> None:
with tempfile.TemporaryDirectory() as directory:
workspace = Path(directory) / "workspace"
user_id = uuid4()
job_id = uuid4()
user_root = workspace / "users" / str(user_id)
working_dir = user_root / "模拟计算"
legacy = user_root / "workspace" / "users" / str(user_id) / "模拟计算" / "origin" / str(job_id)
legacy.mkdir(parents=True)
(legacy / "figure.svg").write_text("<svg/>", encoding="utf-8")
(legacy / "plot-spec.json").write_text("{}", encoding="utf-8")
job = SimpleNamespace(
job_id=job_id,
user_id=user_id,
artifact_manifest=[{
"source_artifact_id": "plot_spec",
"filename": "plot-spec.json",
}],
)
task = SimpleNamespace(
working_dir=f"workspace/users/{user_id}/模拟计算"
)
with patch(
"scripts.repair_software_job_outputs.from_db_path",
return_value=working_dir,
):
plan = build_plan(job, task)
self.assertIsNotNone(plan)
self.assertTrue(plan.move_legacy_dir)
apply_files(plan)
target = working_dir / "origin" / str(job_id)
self.assertTrue((target / "figure.svg").is_file())
self.assertTrue((target / ".meta" / "plot-spec.json").is_file())
self.assertFalse(legacy.exists())
if __name__ == "__main__":
unittest.main()

View File

@ -11,23 +11,24 @@ from alembic.operations import Operations
from sqlalchemy import create_mock_engine
from sqlalchemy.dialects import postgresql
from core.software_jobs import (
_canonical_request,
_published_output_is_valid,
abandon_offer,
list_jobs,
mark_node_jobs_disconnected,
record_job_terminal,
request_job_cancel,
respond_to_offer,
update_job_state,
validate_output_manifest,
)
from core.software_nodes import (
_enrollment_digest,
_hash_secret,
_verify_secret,
delete_node,
)
from core.software_jobs import (
_canonical_request,
abandon_offer,
mark_node_jobs_disconnected,
record_job_terminal,
respond_to_offer,
update_job_state,
validate_output_manifest,
list_jobs,
request_job_cancel,
)
from web.routers.software_nodes import NodeConnectionManager, _bearer
@ -130,8 +131,45 @@ class SoftwareNodeMigrationTests(unittest.TestCase):
self.assertIn("uq_software_jobs_user_idempotency", rendered)
self.assertIn("ix_software_jobs_status_created", rendered)
def test_0033_adds_software_job_artifact_source(self) -> None:
statements: list[str] = []
def capture(sql, *multiparams, **params):
statements.append(str(sql.compile(dialect=postgresql.dialect())))
engine = create_mock_engine("postgresql+psycopg://", capture)
operations = Operations(MigrationContext.configure(engine.connect()))
migration = importlib.import_module(
"db.migrations.versions.20260814_0900_0033_artifact_software_job"
)
with patch.object(migration, "op", operations):
migration.upgrade()
rendered = "\n".join(statements)
self.assertIn("software_job_id", rendered)
self.assertIn("ix_artifacts_software_job_id", rendered)
self.assertIn("jsonb_array_elements", rendered)
class SoftwareJobProtocolTests(unittest.TestCase):
def test_published_output_distinguishes_artifacts_from_metadata(self) -> None:
job_id = uuid4()
self.assertTrue(_published_output_is_valid(job_id, {
"source_artifact_id": "figure_png",
"artifact_id": str(uuid4()),
"path": f"origin/{job_id}/figure.png",
}))
self.assertTrue(_published_output_is_valid(job_id, {
"source_artifact_id": "plot_spec",
"artifact_id": None,
"path": f"origin/{job_id}/.meta/plot-spec.json",
}))
self.assertFalse(_published_output_is_valid(job_id, {
"source_artifact_id": ["plot_spec"],
"artifact_id": None,
"path": f"origin/{job_id}/.meta/plot-spec.json",
}))
@patch("core.software_jobs.session_scope")
def test_queued_job_cancels_without_node_message(self, session_scope) -> None:
session = session_scope.return_value.__enter__.return_value
@ -177,6 +215,7 @@ class SoftwareJobProtocolTests(unittest.TestCase):
self.assertEqual(results[0]["task_name"], "材料仿真")
self.assertEqual(results[0]["node_name"], "LAB-01")
self.assertEqual(results[0]["request_summary"]["display_name"], "Origin 科研绘图")
self.assertEqual(results[0]["output_dir"], f"origin/{job.job_id}")
def test_origin_request_is_canonical_and_rejects_extra_fields(self) -> None:
request = {

View File

@ -20,21 +20,35 @@ class SoftwareOutputPublishTests(unittest.TestCase):
staging.mkdir(parents=True)
working_dir.mkdir()
content = b"origin-result"
metadata = b'{"kind":"plot"}'
(staging / "figure.png").write_bytes(content)
manifest = [{
(staging / "plot-spec.json").write_bytes(metadata)
manifest = [
{
"artifact_id": "figure_png",
"filename": "figure.png",
"media_type": "image/png",
"size_bytes": len(content),
"sha256": hashlib.sha256(content).hexdigest(),
}]
},
{
"artifact_id": "plot_spec",
"filename": "plot-spec.json",
"media_type": "application/json",
"size_bytes": len(metadata),
"sha256": hashlib.sha256(metadata).hexdigest(),
},
]
context = {
"user_id": uuid4(),
"task_id": uuid4(),
"working_dir": "research",
"working_dir": str(working_dir),
}
registered = []
def register(**kwargs):
registered.append(kwargs)
return tuple({
"version": 2,
"scope": "working_dir",
@ -55,10 +69,19 @@ class SoftwareOutputPublishTests(unittest.TestCase):
published = working_dir / "origin" / str(job_id) / "figure.png"
self.assertEqual(published.read_bytes(), content)
self.assertEqual(
(working_dir / "origin" / str(job_id) / ".meta" / "plot-spec.json").read_bytes(),
metadata,
)
self.assertFalse(staging.exists())
self.assertEqual(first[0]["source_artifact_id"], "figure_png")
self.assertEqual(first[0]["path"], f"origin/{job_id}/figure.png")
self.assertEqual(first[1]["source_artifact_id"], "plot_spec")
self.assertIsNone(first[1]["artifact_id"])
self.assertEqual(first[1]["path"], f"origin/{job_id}/.meta/plot-spec.json")
self.assertEqual(second[0]["source_artifact_id"], "figure_png")
self.assertEqual(registered[0]["software_job_id"], job_id)
self.assertEqual(len(registered[0]["refs"]), 1)
if __name__ == "__main__":

View File

@ -12,6 +12,7 @@ from core.software_jobs import (
request_job_cancel,
)
from core.software_nodes import SUPPORTED_CAPABILITIES, list_nodes
from .base import Tool
@ -183,7 +184,11 @@ class SoftwareJobSubmitTool(_SoftwareJobTool):
class SoftwareJobStatusTool(_SoftwareJobTool):
name = "software_job_status"
description = "Check one software job, or list recent jobs in the current task when job_id is omitted."
description = (
"Check one software job, or list recent jobs in the current task when job_id is omitted. "
"A succeeded job returns output_dir as the authoritative starting directory; "
"inspect or search within that directory to analyze its files."
)
parameters = {
"type": "object",
"properties": {"job_id": {"type": "string"}},

View File

@ -9,9 +9,41 @@ from hashlib import sha256
from pathlib import Path
from uuid import UUID
from fastapi import Depends, Header, HTTPException, Request, WebSocket, WebSocketDisconnect, status
from fastapi import (
Depends,
Header,
HTTPException,
Request,
WebSocket,
WebSocketDisconnect,
status,
)
from fastapi.responses import FileResponse
from core.artifact_lifecycle import register_published_artifacts
from core.paths import from_db_path
from core.software_jobs import (
MAX_OUTPUT_ARTIFACT_BYTES,
MAX_OUTPUT_TOTAL_BYTES,
OUTPUT_ARTIFACTS,
SOFTWARE_JOB_METADATA_IDS,
SoftwareJobError,
abandon_offer,
create_job,
get_job,
get_job_input,
get_job_output_context,
list_jobs,
mark_node_jobs_disconnected,
offer_next_job,
pending_node_cancellations,
record_job_terminal,
request_job_cancel,
respond_to_offer,
software_job_output_path,
update_job_state,
validate_output_manifest,
)
from core.software_nodes import (
SoftwareNodeError,
authenticate_node,
@ -23,27 +55,6 @@ from core.software_nodes import (
set_node_disabled,
update_node_runtime,
)
from core.software_jobs import (
MAX_OUTPUT_ARTIFACT_BYTES,
MAX_OUTPUT_TOTAL_BYTES,
OUTPUT_ARTIFACTS,
abandon_offer,
create_job,
get_job,
get_job_input,
get_job_output_context,
mark_node_jobs_disconnected,
offer_next_job,
pending_node_cancellations,
record_job_terminal,
list_jobs,
request_job_cancel,
respond_to_offer,
update_job_state,
validate_output_manifest,
SoftwareJobError,
)
from core.artifact_lifecycle import register_published_artifacts
from web.schemas import (
SoftwareEnrollmentCreateRequest,
SoftwareJobCreateRequest,
@ -153,9 +164,49 @@ def _reject_symlink_path(root: Path, target: Path) -> None:
raise HTTPException(409, "software job output path contains a symbolic link")
def _task_working_dir(root: Path, stored: str) -> Path:
"""解析 tasks.working_dir 的 ROOT 相对 DB 形态,并再次守住 user_root 边界。"""
working_dir = from_db_path(stored).resolve()
try:
working_dir.relative_to(root.resolve())
except ValueError as exc:
raise SoftwareJobError("software job working directory is outside user root") from exc
return working_dir
def _staged_output_path(staging: Path, output_id: str, filename: str) -> Path:
flat = staging / filename
organized = staging / software_job_output_path(output_id)
return flat if flat.is_file() else organized
def _organize_staged_metadata(staging: Path, manifest: list[dict]) -> None:
metadata = [item for item in manifest if item["artifact_id"] in SOFTWARE_JOB_METADATA_IDS]
if not metadata:
return
(staging / ".meta").mkdir(exist_ok=True)
for item in metadata:
source = staging / item["filename"]
destination = staging / software_job_output_path(item["artifact_id"])
if not source.is_file():
if destination.is_file():
continue
raise SoftwareJobError(f"uploaded artifact is missing: {item['artifact_id']}")
if destination.exists():
if (
destination.is_file()
and destination.stat().st_size == item["size_bytes"]
and _hash_file(destination) == item["sha256"]
):
source.unlink()
continue
raise SoftwareJobError("software job metadata destination conflicts")
os.replace(source, destination)
def _publish_software_job_outputs(job_id: UUID, context: dict, manifest: list[dict]) -> list[dict]:
root = load_user_root(context["user_id"])
working_dir = safe_join(root, context["working_dir"])
working_dir = _task_working_dir(root, context["working_dir"])
staging = safe_join(root, f".zcbot_software_job_staging/{job_id}")
relative_output = Path("origin") / str(job_id)
destination = safe_join(working_dir, relative_output.as_posix())
@ -163,7 +214,11 @@ def _publish_software_job_outputs(job_id: UUID, context: dict, manifest: list[di
_reject_symlink_path(root, source)
_reject_symlink_path(root, destination)
for item in manifest:
path = source / item["filename"]
path = (
_staged_output_path(staging, item["artifact_id"], item["filename"])
if source == staging
else destination / software_job_output_path(item["artifact_id"])
)
if (
not path.is_file()
or path.stat().st_size != item["size_bytes"]
@ -171,6 +226,7 @@ def _publish_software_job_outputs(job_id: UUID, context: dict, manifest: list[di
):
raise SoftwareJobError(f"uploaded artifact is missing or invalid: {item['artifact_id']}")
if source == staging:
_organize_staged_metadata(staging, manifest)
destination.parent.mkdir(parents=True, exist_ok=True)
if destination.exists():
raise SoftwareJobError("software job output destination already exists unexpectedly")
@ -180,24 +236,32 @@ def _publish_software_job_outputs(job_id: UUID, context: dict, manifest: list[di
except OSError:
pass
refs = tuple({
"path": (relative_output / item["filename"]).as_posix(),
"path": (relative_output / software_job_output_path(item["artifact_id"])).as_posix(),
"label": item["filename"],
"media_type": item["media_type"],
} for item in manifest)
} for item in manifest if item["artifact_id"] not in SOFTWARE_JOB_METADATA_IDS)
published_refs = register_published_artifacts(
user_id=context["user_id"],
task_id=context["task_id"],
user_root=root,
working_dir=working_dir,
refs=refs,
software_job_id=job_id,
)
refs_by_path = {item["path"]: item for item in published_refs}
return [
{
**item,
"source_artifact_id": item["artifact_id"],
"artifact_id": refs_by_path[(relative_output / item["filename"]).as_posix()]["artifact_id"],
"path": (relative_output / item["filename"]).as_posix(),
"artifact_id": (
refs_by_path.get(
(relative_output / software_job_output_path(item["artifact_id"])).as_posix(),
{},
).get("artifact_id")
),
"path": (
relative_output / software_job_output_path(item["artifact_id"])
).as_posix(),
}
for item in manifest
]
@ -283,9 +347,10 @@ def register_software_node_routes(app, *, require_user, require_admin) -> None:
if len(x_content_sha256) != 64 or any(c not in "0123456789abcdef" for c in x_content_sha256):
raise HTTPException(400, "output artifact digest is invalid")
root = load_user_root(context["user_id"])
working_dir = _task_working_dir(root, context["working_dir"])
published = safe_join(
safe_join(root, context["working_dir"]),
f"origin/{job_id}/{filename}",
working_dir,
f"origin/{job_id}/{software_job_output_path(artifact_id)}",
)
if published.is_file():
if published.stat().st_size == x_content_length and _hash_file(published) == x_content_sha256:
@ -294,14 +359,17 @@ def register_software_node_routes(app, *, require_user, require_admin) -> None:
staging = safe_join(root, f".zcbot_software_job_staging/{job_id}")
_reject_symlink_path(root, staging)
staging.mkdir(parents=True, exist_ok=True)
organized = staging / software_job_output_path(artifact_id)
if organized.is_file():
if organized.stat().st_size == x_content_length and _hash_file(organized) == x_content_sha256:
return None
raise HTTPException(409, "staged output conflicts with uploaded artifact")
destination = staging / filename
if destination.is_file():
if destination.stat().st_size == x_content_length and _hash_file(destination) == x_content_sha256:
return None
raise HTTPException(409, "uploaded output conflicts with existing staging file")
staged_total = sum(
item.stat().st_size for item in staging.iterdir() if item.is_file()
)
staged_total = sum(item.stat().st_size for item in staging.rglob("*") if item.is_file())
if staged_total + x_content_length > MAX_OUTPUT_TOTAL_BYTES:
raise HTTPException(413, "software job outputs exceed the total size limit")
temporary = destination.with_name(destination.name + ".tmp-" + os.urandom(8).hex())