Compare commits

...

2 Commits

Author SHA1 Message Date
caoqianming a1db5756aa feat(tools): 扫描件 PDF 直读 read_document——方舟文档理解 base64 内联(bump 0.58.55)
markitdown 只抽文本层,扫描件(老标准/检测报告/红头指南)转出为空=死路。
探针(scripts/probe_ark_doc.py)验证方舟 chat file 内容块直读 PDF 后落地:

- tools/read_document.py:seed_2_lite 同 variant 同 key,记账走
  record_vision_usage;体积/页数双闸(30MB/100 页,pdfminer 软探页数)+
  finish_reason=length 截断提示;多页 OCR save_md 全文落盘只返 1500 字预览
  (防上下文爆),save_md 与读取侧同款三形态路径解析(相对/宿主绝对/容器
  /workspace)+ user_root 边界
- tools/image_ref.py:抽 load_pdf_as_data_url(复用 resolve_in_root;
  data:application/pdf;base64, 前缀为方舟硬校验)
- config/media/doubao.yaml:vision 段加 max_pdf_mb / max_pdf_pages /
  doc_request_timeout_s(probe 实测口径:~1300 token/页、100 页全覆盖、
  17MB 内联可用、单页栅格化 3600 万像素硬限)
- core/agent_builder.py:注册(跟 look_at_image 同 vision 段走)+ 系统提示
  _MEDIA_READDOC_SEG(何时调/何时不调,防对有文本层 PDF 重复花钱)
- 六 skill(paper/patent/standard/proposal/rebuttal/ppt)摄取段加扫描件
  兜底一行(markitdown 转出空/近空 → read_document OCR)
- 冒烟 scripts/smoke_read_document.py 全过(3 页 ¥0.0066,表格→md 表、
  公式→LaTeX,usage_events 记账验证)

选型对比(不接 MinerU 等外部解析 API=新增第三方数据面;不选本地 OCR=
过度投资;不选 file_url/file_id=免认证公网直链/TOS 依赖)见 DESIGN §8.13。
host 侧工具,无需重建沙箱镜像,pull + 重启 web 生效。

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 16:57:23 +08:00
caoqianming fc53d21620 fix(web): 长对话点目录圆点首次跳不到位——重建窗口瞬时定位+跳转期抑制 sentinel 补载
根因链:loadMessagesAround 后 renderMessages 尾部无条件滚底,把视口钉到
新窗口末尾;底部 sentinel 落在视口内立刻触发 loadNewerMessages,其整窗
重渲染(innerHTML 清空)删掉平滑滚动的目标元素、再把 scrollTop 恢复到
fetch 前位置——首次点击停在窗口底部附近,第二次点击才能到位。

- renderMessages 加 stickBottom 参数(默认 true 保持首屏/发送行为);
  loadEarlier/loadNewer/loadMessagesAround 三个调窗口路径传 false
- jumpToMessage 重建窗口后用瞬时定位(auto),不给补载/图片异步撑高留
  动画窗口期;已在 DOM 的近距离跳转保留平滑
- _msgScrollObserver 在 _outlineJumpLock 期间不补载;解锁时
  _repokeMsgSentinels 对 sentinel unobserve+observe 补投一次交叉状态,
  避免 sentinel 停留在视口内时解锁后无人触发加载

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 16:42:41 +08:00
17 changed files with 757 additions and 10 deletions

View File

@ -5,6 +5,11 @@
> 所以不是每个版本号都有条目。条目格式 `## <版本> — <日期>`,新条目加在最上面。
> 工程口径的完整记录见 `PROGRESS.md` / git log。
## 0.58.55 — 2026-07-21
- 扫描版 PDF 现在也能读了:拍照或扫描生成的 PDF(老标准、检测报告、红头文件等,以前解析出来是空白)会自动逐页识别成文字,表格还原成表格、公式还原成公式,之后写申报书、编标准、做 PPT 都能直接引用其中内容。
- 修复长对话里点右侧目录圆点,第一次点跳不到对应位置、要点第二次才到位的问题。
## 0.58.54 — 2026-07-21
- 手机上打开文件预览的体验优化:预览窗口不再全屏贴边糊满整个屏幕,改为四边留有边距的悬浮卡片,底部消息输入框也不再被挡住,可以边看文件边打字。

View File

@ -341,6 +341,10 @@ scheduled_jobs(§8.5) channel_bindings(§8.7,判别列+JSONB)
- **docker backend**:**专用容器** `zcbot-proc-<id>`(pool.run_proc_container,同款硬化 + iptables init),`product=proc` + 无 instance label —— 与 sandbox 容器的 idle reaper / shutdown_all 生命周期**解耦**,dockerd 托管,蓝绿切换/实例重启不中断。不用 `docker exec -d` 进 sandbox 容器:idle 5min reaper + 启动 shutdown_all 会把长进程随容器带走。
- **回收**:check_process 见终态顺手 rm 容器;web lifespan 每小时 `procs.sweep`(终态目录 7d TTL / exited 孤儿容器),幂等,蓝绿双实例同时跑无害。
- **通知/可视**:不做服务端推送 —— 前端轮询 `GET /v1/procs`(用户级,纯文件读取,仅有 running proc 时 5s 一拉):`[Background]` 工具结果卡本身活化(spinner+跳秒+停止按钮,与前台工具卡同体验,历史重渲同样恢复;`POST .../procs/<id>/kill`)、running→终态弹 toast(跨 task 也提醒,点击跳转)。proc 完成时刻往往没有活跃 run,SSE 通道根本不在,轮询是诚实的选型。
### 8.13 扫描件 PDF 直读:方舟文档理解,不接外部 OCR(✅ 2026-07-21)
缺口:markitdown 只抽 PDF 文本层,扫描件(老标准/检测报告/红头指南,建材院高频)转出为空=死路。**选复用 seed-2.0-lite 的方舟文档理解**(chat file 内容块,PDF 整本 base64 内联)新增 `read_document`:零新供应商(敏感文档不出已有豆包面)、零新基础设施、记账复用 vision 通道;实测 ~1300 输入 token/页(约 1 厘/页)、100 页全覆盖、17MB 内联可用。**不选专用解析 API**(MinerU/Textin:版面还原最好,但申报书/专利底稿要上传新第三方 + 免费额度政策不稳);**不选本地 OCR**(PaddleOCR 类:镜像塞推理依赖,需求未量化前过度投资);**不选 file_url/file_id 传址**(前者要给用户文件开免认证公网直链=新安全面、开发机 NAT 后还跑不通;后者要接 TOS 多落一份存储;base64 是零新增面的唯一形态,行业惯例 chat 端点也不收 multipart)。防上下文爆:多页 OCR 强制 `save_md` 落盘只返预览。**升级信号**:>100 页/>30MB 巨件成高频 → 接 TOS 走 file_id;要高保真版面/公式还原 → 再评 MinerU。probe/smoke 留仓(`scripts/probe_ark_doc.py` / `smoke_read_document.py`)。
- **对话锁(前端)**:bg proc 运行期间该 task 的 composer 锁定(发送→停止,Enter 拦截),观感与前台执行完全一致 —— 后台化的收益定位为「进程扛超时/服务重启」,**不改变"一个任务同时只做一件事"的对话心智**;完成的那次轮询解锁 + toast「可继续对话」。锁只在前端,服务端不 409:「停止」入口必须可达,且多设备/渠道绕过前端锁属可接受边缘(等的是同一个进程,发了消息也不冲突)。
- **防失控**:每用户并发 running 上限(`ZCBOT_MAX_BG_PROCS` 默 3);前台默认超时不放大(它是逼模型做前台/后台选择的杠杆)。

View File

@ -2,7 +2,7 @@
> 配合 `DESIGN.md`。本文件只记 phase 状态、决策偏差、文件量、下一步。每条 1-2 句:做了啥 + 关键判断;细节查 `git log` / `git diff` / `DESIGN §7.9`
最后更新:2026-07-21(系统提示加护栏:大段中文正文别内联进 .py——run_python 语法预检失败第一大头治本,bump 0.58.51)
最后更新:2026-07-21(扫描件 PDF 直读 read_document:方舟文档理解 base64 内联,markitdown 死路补位,bump 0.58.55)
---
@ -23,6 +23,8 @@
### 2026-07
- **07-21 / 0.58.55 / 扫描件 PDF 直读(read_document,方舟文档理解)**:markitdown 只抽文本层,扫描件(老标准/检测报告/红头指南)转出为空=死路。探针(`scripts/probe_ark_doc.py`)验证方舟 chat file 内容块直读 PDF:格式 `{"type":"file","file":{filename,file_data}}` + `data:application/pdf;base64,` 前缀、base64 内联 17MB 可用(免 TOS)、单页栅格化 3600 万像素硬限(PIL 存 PDF 需标对 dpi)、~1300 输入 token/页(约 1 厘/页)、100 页魔术串全覆盖。落地 `tools/read_document.py`(seed_2_lite 同 variant 同 key,记账走 record_vision_usage):体积/页数双闸(30MB/100 页,pdfminer 软探页数免白付撞上下文)+ `finish_reason=length` 截断提示 + **多页 OCR `save_md` 全文落盘只返 1500 字预览**(防上下文爆);`image_ref.py` 抽 `load_pdf_as_data_url` 复用三形态路径解析与 user_root 边界。agent_builder 注册 + 系统提示 `_MEDIA_READDOC_SEG`(何时调/何时不调防重复花钱);六 skill(paper/patent/standard/proposal/rebuttal/ppt)摄取段加扫描件兜底一行。冒烟 `scripts/smoke_read_document.py` 全过(3 页 ¥0.0066,表格→md 表、公式→LaTeX)。选型对比(外部解析 API=新增第三方数据面 / 本地 OCR=过度投资)见 DESIGN §8.13。host 侧工具,**无需重建沙箱镜像**。
- **07-21 / 0.58.55 / 长对话点目录圆点首次跳不到位修复**:根因链=loadMessagesAround 后 renderMessages 尾部无条件滚底钉到窗口末尾,底部 sentinel 入视口立刻触发 loadNewerMessages 整窗重渲染删掉平滑滚动目标。修:renderMessages 加 stickBottom 参数(三个调窗口路径传 false);jumpToMessage 重建窗口后瞬时定位(auto)不留动画窗口期;`_msgScrollObserver` 在 `_outlineJumpLock` 期间不补载、解锁时对 sentinel 重投交叉状态。
- **07-21 / 0.58.54 / 手机端文件预览改悬浮卡片(四边留边距、不压输入区)**:用户反馈移动端预览弹框全屏贴边(100vw×100dvh、直角)观感差且糊在消息输入框上。改 `dev.html` ≤640px 媒体查询:遮罩层 `padding:10px` 四边留呼吸边距、底部 `calc(--preview-bottom-inset + 10px)` 在 chat-form 让位之上再加间隙;卡片尺寸改 `100%`(相对遮罩内容区,弃 100dvh,旧 WebView 兼容更稳),恢复默认圆角。输入区仍走既有 `body.fp-open #chat-form` z-index 抬升保持可用;桌面端不动。
- **07-21 / 0.58.53 / App 套壳进入契约(embed app 变体 relogin_url)+ 对接文档**:移动 App 方案定型——原生壳(WebView)+ 原生登录页,登录走 platform 自有接口(`/api/login/token` → `/api/login/external-login`,已实测打通,后者服务端用 PLATFORM_KEY 换 zcbot JWT,user_id=platform user.uuid)。顶层 WebView 无父窗口,iframe 的 postMessage 协议失效(`zcbot-401` 经 `window.parent` 发不出去),故把企微免登模式泛化为 **app 变体**:`?embed=1&relogin_url=<绝对地址>#token=..&user_id=..` —— fragment 注入(读完即清,同 wecom)+ 401/logout 时 `location.replace(relogin_url)`(原生壳拦自定义 scheme 如 `zcbotapp://relogin` → 静默换新 token 重进;H5 登录页同契约)。改动:`state.js` 加 `EMBED_RELOGIN_URL` 解析+消毒(必须带 scheme 的绝对地址,拦 `javascript:`/`data:` 等可执行 scheme——该值喂 `location.replace`,不消毒是 XSS 口子);`embed.js` 抽共用 `readFragmentToken`/`gotoInitialTask`,加 `embedAppInit`/`embedAppRelogin`;`auth.js` logout 分支序 wecom→app→iframe。新增 `APP.md`(进入契约 + platform 登录链路实测 + 原生壳杂活清单 + H5 备选),`EMBED.md` 精简 243→约 120 行。ESM 语法检查过,浏览器级实测待跑(逻辑与线上企微免登同构)。顺带:实测发现 platform `/api/login/token` 响应把 `hashed_password` 下发到客户端,修复建议已写进 APP.md §5 转交对方。
- **07-21 / 0.58.52 / 窗口体量估算实测校准(50%压缩/85%折叠/前端占用环共用)**:diag 复盘窗口约束现状(`scripts/diag_context_pressure.py` 留仓):机制有效(折叠上线后唯一越线 task 9a863424 下个 run 起点即回落 400k→137k)、零撞硬上限(6 条终态错误无一 context 超限),但量出静态 `CHARS_PER_TOKEN=2.5` 对中文密集窗口**低估近一倍**——名义 85% 折叠线实际 ~155% reliable 才触发(该 task 实测 40.7 万 tokens);代码密集反向虚高(fe2d8b73 估 1.2M 实际 616k)。**修(信号校准,不加新机制)**:`context.py` 加 `estimate_window_tokens`(provider 实报 tokens_in/out 覆盖窗口主体,仅实测点后尾巴按 2.5 估)+ `calibrated_chars_per_token`(比值夹 [1.0,4.0] 带宽);`Session.last_measured_usage()` 从 messages 表取窗口内最后一条实报 usage(best-effort 绝不抛,idx→内存 pos 映射校验 role);`maybe_fold` 触发判定换 token 实测口径;loop 的压缩门槛与 `context_limit_chars`(前端环)用校准比值,每轮成功调用后以 (sent_chars/prompt_tokens) 刷新。已知残余:折叠后 run 若在首次 chat 完成前崩,下个 run 起点读到旧实测会多折一次(后果=摘要偏保守,原文全在 DB,不为此加持久化状态)。run 中途折叠/超限自愈按 §5 无信号不实施继续搁置。真实生产 task 验证映射与校准(9a863424 drift 0.95x、fe2d8b73 0.51x、74696048 0.86x),286 测试全绿。
@ -231,7 +233,7 @@ core/asr_xfyun.py 170 ← 讯飞语音听写 IAT wss 客户端(整段
core/asr_lfasr.py 250 ← 讯飞录音文件转写 LFASR 客户端(异步订单 + 说话人分离;transcribe_audio 工具底座,diag: scripts/diag_lfasr.py)
core/agent_builder.py 340 ← 装配 lib(有 ARK_API_KEY 才挂 SeedreamTool);build_skill_registry 装两来源
core/executor.py / sandbox/{network,pool}.py / executor_docker.py ← Executor ABC + Docker per-user 容器池
tools/{base,fs,shell,run_python,skill_tool,skill_authoring,seedream,seedance,web_search,web_fetch,documents,materials_project,transcribe_audio}.py ← skill_authoring=save_skill/fork_skill(host-side 写 user .skills)
tools/{base,fs,shell,run_python,skill_tool,skill_authoring,seedream,seedance,look_at_image,read_document,image_ref,web_search,web_fetch,documents,materials_project,transcribe_audio}.py ← read_document=扫描件 PDF OCR(方舟文档理解);image_ref=图/PDF 路径解析+base64 共享
main.py ~210 ← 入口:web / db / probe / user / sandbox check
db/migrations/versions/ 0001-0008
web/app.py ~1360 ← /v1 JSON API + user_id 隔离 + run lock + cancel + files + pptx 预览 + skills(列表/正文/删)

View File

@ -47,6 +47,12 @@ vision:
max_image_mb: 10 # 单图上限(超出 tool 侧直接报错,不发请求)
request_timeout_s: 120 # 读图慢于此判超时(非流式,长 OCR 首字节可能逼近上限)
timeout_retries: 1 # 超时/网络抖动 tool 内透明重试次数(退避 2^n s);不含业务错误
# ---- 文档理解(read_document,同模型同价,scripts/probe_ark_doc.py 实测口径) ----
# 方舟 file 内容块直读 PDF(含扫描件):base64 内联,约 1300 token/页(每页约 1 厘)。
# 单页栅格化 3600 万像素硬限在 ARK 侧;这里只闸文件体积与页数(页数≈上下文占用)。
max_pdf_mb: 30 # 单 PDF 体积上限(base64 后 x1.33;实测 17MB 内联可用)
max_pdf_pages: 100 # 单次页数上限(100 页≈13 万输入 token);更长先拆分卷
doc_request_timeout_s: 600 # 整本 OCR 输出量大,显著慢于单图(非流式)
video:
# fast 放第一个 → 默认 variant(成本敏感场景优先);开通了 Pro 的用户从顶栏下拉切。

View File

@ -1,3 +1,3 @@
# zcbot 版本号单一事实源:web/app.py 的 FastAPI version、/healthz 返回、前端展示都引这里。
# 改版本只动这一行。
__version__ = "0.58.54"
__version__ = "0.58.55"

View File

@ -46,6 +46,7 @@ from tools.materials_project import (
MaterialsProjectSearchSummaryTool,
)
from tools.look_at_image import LookAtImageTool
from tools.read_document import ReadDocumentTool
from tools.check_process import CheckProcessTool
from tools.run_python import RunPythonTool
from tools.seedance import SeedanceTool
@ -78,6 +79,10 @@ _MEDIA_LOOK_SEG = """\
- `look_at_image` 看图 / 读图(豆包 Seed 2.0 Lite 视觉)**(主模型)是纯文本看不见图,""图就调它**:OCR 文字描述画面读图表/表格/示意图识别物体**很慢(一次要几十秒),谨慎调用** 只在确实需要图的实际内容才调
- **何时调**:用户消息里出现 `[用户上传的参考图] <路径>` 且需要据图内容回答("这图里写了啥 / 是什么 / 表格数据多少");或要基于 task 内某张图(`figures/xxx.png`)**实际内容**做事(不是改图,改图走 seedream) `image=<路径>` + 可选 `question`,**把想知道的一次在 `question` 里问全**,别分多次看
- **何时不调**:用户只是要改图( seedream i2i)/ 只要文件名不关心内容 / 图是你自己刚生成的且 prompt 已知(无需再读)/ 图的内容对当前任务可有可无**绝不对同一张图反复看**(每看一次都是几十秒等待)"""
_MEDIA_READDOC_SEG = """\
- `read_document` PDF(豆包 Seed 2.0 Lite 文档理解),**专治扫描件**:markitdown 对某 PDF 转出** / 近空**(纯图页无文本层) 用它逐页 OCR markdown每页约 1-2 厘钱,单次上限 100 (更长先拆分卷)
- **多页整本 OCR 必传 `save_md`**( `save_md='source/xxx.md'`):全文落文件只返回预览,后续直接 read 那个 md;只问局部内容("第3章指标是什么") `question` 即可不落盘
- **何时不调**:PDF 有文本层且 markitdown 已转出正文(直接用那份 md,别重复花钱);DOCX/PPTX/XLSX( markitdown);单张图片( look_at_image)"""
_MEDIA_SEEDREAM_SEG = """\
- `seedream` 豆包图像生成 / 改图产物自动落 `<task_dir>/figures/`每次 **¥0.22**(联网 `search=true` ¥0.05)
- **文生图**(不传 `reference_images`):从零按 prompt **改图 i2i**( `reference_images=["figures/xxx.png"]`):在已有图上做像素级修改**用户对刚生成 / 上传的图说"改成 X / 换个颜色 / 去掉某处" 必须走改图(reference_images 指那张图),绝不重新文生图**(重画 = 完全不同的图,丢原构图)v1 改图仅支持单张参考
@ -105,6 +110,7 @@ def _media_tools_block(has_ark: bool, image_tool: str) -> str:
segs: list[str] = []
if has_ark:
segs.append(_MEDIA_LOOK_SEG)
segs.append(_MEDIA_READDOC_SEG)
if image_tool == "seedream":
segs.append(_MEDIA_SEEDREAM_SEG)
segs.append(_MEDIA_DIAGRAM_FORK_SEG)
@ -763,6 +769,18 @@ def build_agent(
user_root=ur_path,
)
tools[look_tool.name] = look_tool
# 文档理解(read_document):同 variant 同 key,扫描件 PDF OCR(markitdown 死路补位)
readdoc_tool = ReadDocumentTool(
ark_cfg=ark_cfg,
vision_variant_cfg=vis_variant,
variant_key=vis_key,
working_dir=working_dir_path,
task_id=task_id,
user_id=uid,
base_dir=tool_base,
user_root=ur_path,
)
tools[readdoc_tool.name] = readdoc_tool
# 录音文件转写(transcribe_audio / 讯飞 LFASR):仅当 XFYUN_APPID +
# XFYUN_LFASR_SECRET_KEY 齐了才挂(沿用"有 key 才注册")。与 IAT 语音听写是两个

248
scripts/probe_ark_doc.py Normal file
View File

@ -0,0 +1,248 @@
"""Probe: 方舟"文档理解"能不能直读扫描件 PDF(路线 A 验证,不动线上代码)。
跑法: .venv/Scripts/python.exe scripts/probe_ark_doc.py [真实PDF路径]
依赖 .env ARK_API_KEY**会真调豆包 seed-2.0-lite,产生 < ¥0.05 费用**
验证点:
1. 合成一份 3 "扫描件"PDF(PIL 纯图页,无文本层,每页埋魔术串 + 中文段落/表格)
2. markitdown 对它抽不出文字(证明现状确实是死路)
3. chat/completions file 内容块用哪种 JSON 形状能被接受(候选格式挨个试,
400 报错会带字段名,本身就是探针产出)
4. OCR 保真度:三页魔术串是否全命中(页覆盖)中文/表格内容是否读出
5. usage tokens 每页成本口径
传真实 PDF 路径时跳过合成,直接测该文件(页数/体积上限真实扫描件质量)
"""
from __future__ import annotations
import base64
import json
import os
import subprocess
import sys
import tempfile
from pathlib import Path
ROOT = Path(__file__).resolve().parent.parent
sys.path.insert(0, str(ROOT))
# Windows 控制台默认 GBK,打印中文/特殊符号会崩 → 强制 stdout UTF-8
try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace") # type: ignore[attr-defined]
except Exception:
pass
# 读 .env(同 smoke_look_at_image)
env_file = ROOT / ".env"
if env_file.exists():
for line in env_file.read_text(encoding="utf-8").splitlines():
line = line.strip()
if not line or line.startswith("#") or "=" not in line:
continue
k, _, v = line.partition("=")
os.environ.setdefault(k.strip(), v.strip())
from PIL import Image, ImageDraw, ImageFont
from core.ark_client import ArkClient, ArkConfig, ArkError
def _magic(i: int) -> str:
"""第 i 页(0 起)的魔术串,确定性可复算。"""
return f"ZCBOT-DOC-P{i + 1}-{7391 + i * 613}"
# 每页正文:标题 + 段落 + 第 2 页一个小表格,模拟真实扫描标准的版面
PAGE_LINES = [
[
"水泥胶砂强度检验方法(模拟扫描件 第1页)",
"本方法规定了水泥胶砂抗压强度与抗折强度的测定步骤。",
"试验室温度应保持在 20 ± 2 摄氏度,相对湿度不低于 50%",
"校验码: {magic}",
],
[
"第2页 配合比与龄期",
"胶砂配比 水泥 450g 标准砂 1350g 水 225g",
"龄期(d) 3 7 28",
"抗压(MPa) 22.5 35.0 52.5",
"校验码: {magic}",
],
[
"第3页 结果处理",
"以三条试体抗折结果的算术平均值作为试验结果。",
"当三个值中有超出平均值 ±10% 时,应剔除后重新计算。",
"校验码: {magic}",
],
]
_FONT_CANDIDATES = [
r"C:\Windows\Fonts\msyh.ttc",
r"C:\Windows\Fonts\simhei.ttf",
r"C:\Windows\Fonts\simsun.ttc",
]
def _load_font(size: int) -> ImageFont.FreeTypeFont | ImageFont.ImageFont:
for p in _FONT_CANDIDATES:
if Path(p).exists():
return ImageFont.truetype(p, size)
return ImageFont.load_default()
def make_scanned_pdf(dest: Path, n_pages: int = 3) -> None:
"""N 页纯图 PDF(A4 150dpi,无文本层)—— 模拟扫描件。
3 页用真实版面(段落 + 表格),之后的页是短页(标题 + 魔术串),
专测页数上限 / 页覆盖1-bit 模式控体积(黑字白底扫描件本就近似双色)
"""
font = _load_font(36)
pages = []
for i in range(n_pages):
img = Image.new("RGB", (1240, 1754), (255, 255, 255))
d = ImageDraw.Draw(img)
lines = (PAGE_LINES[i] if i < len(PAGE_LINES)
else [f"{i + 1}页 附录条款", "本页为附录占位内容。", "校验码: {magic}"])
y = 120
for line in lines:
d.text((100, y), line.format(magic=_magic(i)), fill=(0, 0, 0), font=font)
y += 90
# 1-bit 模式绕开本环境 PIL 缺 JPEG 编码器的问题(RGB 页会走 DCT/JPEG),
# 且体积最小(P 模式实测一页 4MB+,多页测试撑爆请求)
pages.append(img.convert("1", dither=Image.Dither.NONE))
# resolution=150:页物理尺寸=px/150 英寸 ≈ A4。缺省 72dpi 会把页标成 A4 两倍大,
# 方舟按固定 dpi 栅格化 PDF 页,超尺寸页撞"单页 3600 万像素"上限(probe 实测报错)
pages[0].save(dest, save_all=True, append_images=pages[1:], resolution=150.0)
def check_markitdown_dead_end(pdf: Path) -> None:
"""现状对照:markitdown 对纯图 PDF 应抽不出正文。"""
exe = ROOT / ".venv" / "Scripts" / "markitdown.exe"
cmd = [str(exe) if exe.exists() else "markitdown", str(pdf)]
try:
r = subprocess.run(cmd, capture_output=True, text=True, timeout=120,
encoding="utf-8", errors="replace")
text = (r.stdout or "").strip()
print(f"[markitdown] exit={r.returncode} 抽出正文 {len(text)} 字符"
+ (f" → 非空?! 前 200 字: {text[:200]!r}" if text else " → 空(证实扫描件死路)"))
except Exception as e:
print(f"[markitdown] 跑不了({type(e).__name__}: {e}),跳过对照")
# file 内容块候选形状:方舟文档说 file_id/file_data/file_url 三选一,但块的外层
# JSON 没抓到 → 挨个试,400 报错信息(带字段名)也是探针产出
def _candidate_blocks(b64: str, filename: str) -> list[tuple[str, dict]]:
data_url = f"data:application/pdf;base64,{b64}"
return [
("openai_file_data_url",
{"type": "file", "file": {"filename": filename, "file_data": data_url}}),
("openai_file_raw_b64",
{"type": "file", "file": {"filename": filename, "file_data": b64}}),
("ark_file_url_data",
{"type": "file_url", "file_url": {"url": data_url}}),
]
QUESTION = (
"这是一份多页 PDF 文档。请逐页把其中的文字完整 OCR 出来,"
"每页以「== 第N页 ==」开头,保留表格数据与换行,不要总结不要遗漏。"
)
# 多页模式只要校验码清单:把"页覆盖上限"和"输出 token 上限"两个变量拆开测
QUESTION_MAGIC_ONLY = (
"这是一份多页 PDF 文档,每页都有一行「校验码: ZCBOT-DOC-...」。"
"请按页序把每页的校验码逐行列出(格式:第N页 <校验码>),只要校验码,别的不用输出。"
)
def main() -> int:
cfg = ArkConfig.load()
if cfg is None:
print("[SKIP] ARK_API_KEY 未设(或 doubao.yaml 缺失)")
return 0
vcfg = (cfg.raw.get("vision") or {}).get("seed_2_lite") or {}
model_id = vcfg.get("model_id", "doubao-seed-2-0-lite-260428")
print(f"[setup] model={model_id} base={cfg.base_url}")
n_pages = 3
if len(sys.argv) > 2 and sys.argv[1] == "--pages":
n_pages = int(sys.argv[2])
pdf = None
elif len(sys.argv) > 1:
pdf = Path(sys.argv[1])
else:
pdf = None
if pdf is not None:
synthetic = False
print(f"[setup] 使用真实 PDF: {pdf}")
else:
pdf = Path(tempfile.mkdtemp(prefix="zcbot_probe_")) / "scanned_probe.pdf"
make_scanned_pdf(pdf, n_pages)
synthetic = True
print(f"[setup] 合成扫描件 PDF: {pdf}"
f"({pdf.stat().st_size} bytes, {n_pages} 页, 无文本层)")
check_markitdown_dead_end(pdf)
b64 = base64.b64encode(pdf.read_bytes()).decode()
print(f"[setup] base64 体积 {len(b64) / 1024:.0f} KB")
question = QUESTION_MAGIC_ONLY if (synthetic and n_pages > 5) else QUESTION
resp = None
accepted = None
for name, block in _candidate_blocks(b64, pdf.name):
body = {
"model": model_id,
"messages": [{
"role": "user",
"content": [{"type": "text", "text": question}, block],
}],
}
print(f"[try] 格式 {name} ...")
try:
with ArkClient(cfg, timeout_s=300) as client:
resp = client.post_json("/chat/completions", body, timeout_s=300)
accepted = name
print(f"[OK] 格式 {name} 被接受")
break
except ArkError as e:
print(f"[reject] {name}: {e}")
if resp is None:
print("\n[FAIL] 所有候选格式都被拒 —— 看上面报错定位正确字段名,"
"或该模型版本不支持 file 输入(需查文档理解模型列表)")
return 2
content = ((resp.get("choices") or [{}])[0].get("message") or {}).get("content") or ""
if isinstance(content, list):
content = "\n".join(c.get("text", "") for c in content if isinstance(c, dict))
usage = resp.get("usage") or {}
tin = int(usage.get("prompt_tokens", 0) or 0)
tout = int(usage.get("completion_tokens", 0) or 0)
cost = (tin * float(vcfg.get("price_cny_per_mtoken_input", 0.6))
+ tout * float(vcfg.get("price_cny_per_mtoken_output", 3.6))) / 1e6
print(f"\n[usage] tokens={tin}+{tout} cost≈¥{cost:.4f} accepted_format={accepted}")
print(f"[response]\n{content}\n")
if synthetic:
flat = content.replace(" ", "").replace("-", "")
magics = [_magic(i) for i in range(n_pages)]
hits = [m for m in magics if m.replace("-", "") in flat]
missed = [m for m in magics if m not in hits]
print(f"[verify] 魔术串命中 {len(hits)}/{n_pages}"
+ (f" 漏: {missed[:5]}{'...' if len(missed) > 5 else ''}" if missed else ""))
if n_pages <= 5:
cn_hit = "标准砂" in content and "抗压" in content
print(f"[verify] 中文表格关键词(标准砂/抗压)命中: {cn_hit}")
else:
cn_hit = True
if not missed and cn_hit:
print(f"\n[PASS] {n_pages} 页全覆盖:file 输入 + 扫描件 OCR 验证通过")
return 0
print("\n[WARN] 部分未命中 —— 人工核对上面 response 判断保真度")
return 1
print("[DONE] 真实 PDF 模式:人工核对上面 response")
return 0
if __name__ == "__main__":
sys.exit(main())

View File

@ -0,0 +1,116 @@
"""Smoke: read_document(豆包文档理解)端到端走通 + 扫描件 OCR + save_md 落盘验证。
跑法: .venv/Scripts/python.exe scripts/smoke_read_document.py
依赖 .env ARK_API_KEY / ZCBOT_DB_URL**会真调豆包文档理解,产生 < ¥0.01 费用**
校验:
1. 合成 3 页扫描件 PDF(复用 probe_ark_doc 的生成器,无文本层)
2. ReadDocumentTool.execute(save_md=...) 返回 banner + saved: + 预览
3. save_md 文件落盘且三页魔术串全命中(页覆盖 + OCR 保真)
4. usage_events 多一行 kind="vision",units document 路径
5. _count_pdf_pages 软探页数 = 3(页数闸门的数据源)
"""
from __future__ import annotations
import sys
import uuid
from pathlib import Path
ROOT = Path(__file__).resolve().parent.parent
sys.path.insert(0, str(ROOT))
sys.path.insert(0, str(ROOT / "scripts"))
try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace") # type: ignore[attr-defined]
except Exception:
pass
# probe_ark_doc import 时自带 .env 加载 + sys.path 处理
from probe_ark_doc import _magic, make_scanned_pdf
from sqlalchemy import text
from core.ark_client import ArkConfig
from core.storage import session_scope
from core.storage.models import Task, User
from tools.read_document import ReadDocumentTool, _count_pdf_pages
def main() -> int:
cfg = ArkConfig.load()
if cfg is None:
print("[SKIP] ARK_API_KEY 未设(或 doubao.yaml 缺失),无法测真接口")
return 0
vision_cfg = (cfg.raw.get("vision") or {})
if not vision_cfg:
print("[SKIP] doubao.yaml 无 vision 段")
return 0
variant_key, variant_cfg = next(iter(vision_cfg.items()))
print(f"[setup] variant={variant_key} model={variant_cfg.get('model_id')} "
f"max_pdf_mb={variant_cfg.get('max_pdf_mb')} "
f"max_pdf_pages={variant_cfg.get('max_pdf_pages')}")
uid = uuid.uuid4()
tid = uuid.uuid4()
ws_user = ROOT / "workspace" / "users" / str(uid)
wd = ws_user / "smoke_readdoc"
pdf = wd / "upload" / "scan.pdf"
pdf.parent.mkdir(parents=True, exist_ok=True)
make_scanned_pdf(pdf, 3)
print(f"[setup] 合成扫描件 {pdf.name}({pdf.stat().st_size} bytes, 3 页, 无文本层)")
n = _count_pdf_pages(pdf)
assert n == 3, f"_count_pdf_pages 应 3,实际 {n}"
print(f"[OK] _count_pdf_pages = {n}")
with session_scope() as s:
s.add(User(user_id=uid))
with session_scope() as s:
s.add(Task(task_id=tid, user_id=uid, name="smoke_readdoc", working_dir=str(wd)))
tool = ReadDocumentTool(
ark_cfg=cfg,
vision_variant_cfg=variant_cfg,
variant_key=variant_key,
working_dir=wd,
task_id=tid,
user_id=uid,
base_dir=wd,
user_root=ws_user,
)
print("[call] execute(document='upload/scan.pdf', save_md='source/scan.md')")
result = tool.execute(document="upload/scan.pdf", save_md="source/scan.md")
print(f"[tool result]\n{result}\n")
if result.startswith("[Error]"):
print("[FAIL] tool 返回错误")
return 2
assert "saved:" in result, "返回里缺 saved: 行"
saved = wd / "source" / "scan.md"
assert saved.is_file(), f"save_md 未落盘: {saved}"
body = saved.read_text(encoding="utf-8")
flat = body.replace(" ", "").replace("-", "")
magics = [_magic(i) for i in range(3)]
hits = [m for m in magics if m.replace("-", "") in flat]
assert len(hits) == 3, f"魔术串命中 {len(hits)}/3(漏 {set(magics) - set(hits)})"
print(f"[OK] save_md 落盘 {len(body)} 字符,3/3 魔术串命中")
with session_scope() as s:
rows = s.execute(text(
"SELECT kind, model_profile, units, cost_cny FROM usage_events "
"WHERE task_id = :tid"
), {"tid": str(tid)}).all()
assert len(rows) == 1, f"usage_events 行数应 1,实际 {len(rows)}"
row = rows[0]
assert row.kind == "vision", f"kind 应 vision,实际 {row.kind}"
assert "document" in row.units, f"units 缺 document: {row.units}"
print(f"[OK] usage_events: kind={row.kind} model={row.model_profile} "
f"cost_cny={row.cost_cny} units={row.units}")
print("\n[PASS] smoke_read_document 全部通过")
return 0
if __name__ == "__main__":
sys.exit(main())

View File

@ -62,6 +62,8 @@ markitdown <path>/ref_paper.pdf -o <task_dir>/source/ref.md
markitdown https://.../guide -o <task_dir>/source/guide.md
```
某 PDF 转出为空 / 近空 = 扫描件(纯图无文本层)→ 改用 `read_document` 工具 OCR:`read_document(document='<path>/xxx.pdf', save_md='source/xxx.md')`,产物同样是 md。
转完后阶段一直接 `read <task_dir>/source/*.md` 拿事实,**实验数据一律以用户素材为准,不得自造**。
## 阶段一:八条对齐(写 spec)

View File

@ -30,6 +30,8 @@ markitdown <path>/论文.pdf -o <task_dir>/source/论文.md
markitdown https://example.com/ -o <task_dir>/source/外部.md
```
某 PDF 转出为空 / 近空 = 扫描件(纯图无文本层)→ 改用 `read_document` 工具 OCR:`read_document(document='<path>/xxx.pdf', save_md='source/xxx.md')`,产物同样是 md。
代码仓库 / 单文件代码 → 直接 `read`,关键算法在哪个函数、参数怎么传、跟现有方案差异在哪 —— 边读边记到 spec 草稿里。
## 阶段一: 专利点挖掘与筛选

View File

@ -38,7 +38,7 @@ model: glm.pro52
- `icons/` —— **5 套图标库**(tabler-outline/tabler-filled/chunk-filled/phosphor-duotone/simple-icons,共 1.1w+)。executor 写 `<use data-icon="<lib>/<name>">`,finalize 自动从这里内嵌(默认目录,无需预取);锁 inventory 前用 `ls templates/icons/<lib>/ | grep <关键词>` 验名
- `design_spec_reference.md` / `spec_lock_reference.md` —— **spec 产出骨架**,策略阶段写 spec 前必读
**素材摄取**:用 `markitdown` CLI 把 PDF/DOCX/PPTX/XLSX/HTML/URL 转 Markdown,落 `<project_dir>/sources/<name>.md`
**素材摄取**:用 `markitdown` CLI 把 PDF/DOCX/PPTX/XLSX/HTML/URL 转 Markdown,落 `<project_dir>/sources/<name>.md`某 PDF 转出为空 / 近空 = 扫描件(纯图无文本层)→ 改用 `read_document` 工具 OCR(`save_md` 指到同一位置),产物同样是 md。
## 工作目录约定

View File

@ -35,6 +35,8 @@ markitdown <path>/budget.xlsx -o <task_dir>/source/budget.md
markitdown https://example.com/x -o <task_dir>/source/policy.md
```
某 PDF 转出为空 / 近空 = 扫描件(红头指南常见,纯图无文本层)→ 改用 `read_document` 工具 OCR:`read_document(document='<path>/xxx.pdf', save_md='source/xxx.md')`,产物同样是 md。
转完后 spec 阶段直接 `read <task_dir>/source/*.md` 拿事实,不要凭印象写。
## 阶段一: 八条对齐

View File

@ -48,6 +48,8 @@ markitdown <path>/decision_letter.eml -o <task_dir>/source/decision.md
markitdown <path>/manuscript.docx -o <task_dir>/source/manuscript.md
```
某 PDF 转出为空 / 近空 = 扫描件(纯图无文本层)→ 改用 `read_document` 工具 OCR:`read_document(document='<path>/xxx.pdf', save_md='source/xxx.md')`,产物同样是 md。
缺原稿也能做分诊和回复框架,但**改稿说明只能占位**;主动向用户要。用户若有自己的逐条应对腹稿(中文随手记也行),一并收进 `source/notes.md` —— 这是后面动作映射的关键输入。
## 阶段一:拆条 + 对齐 spec

View File

@ -39,6 +39,8 @@ markitdown <path>/数据.xlsx -o <task_dir>/source/data.md
markitdown https://example.com/ -o <task_dir>/source/policy.md
```
某 PDF 转出为空 / 近空 = 扫描件(老标准 / 检测报告常见,纯图无文本层)→ 改用 `read_document` 工具 OCR:`read_document(document='<path>/xxx.pdf', save_md='source/xxx.md')`,产物同样是 md。
## 阶段一: 定层级与体裁 + 八条对齐 (spec)
产物:**task 级 spec 文件**(标准"宪法",阶段二每章前都要重读)。文件路径按 system prompt 的《task 级「宪法」文件命名约定》:

View File

@ -104,3 +104,41 @@ def load_image_as_data_url(
b64 = base64.b64encode(raw).decode("ascii")
return f"data:{mime};base64,{b64}", display_fn(resolved), ""
def load_pdf_as_data_url(
rel: str,
*,
working_dir: Path,
user_root: Optional[Path],
display_fn: Callable[[Path], str],
max_bytes: int,
) -> tuple[str, str, str]:
"""PDF 版 load_image_as_data_url:同一套三形态解析 + user_root 边界 + 大小上限。
返回 (data_url, display_path, error);data URL 前缀必须是
`data:application/pdf;base64,`(方舟硬校验,probe_ark_doc 实测)
"""
resolved = resolve_in_root(rel, working_dir, user_root)
if resolved is None:
return "", "", (
f"[Error] 文档找不到或越界: {rel!r}。请传 task_dir 内已存在 PDF 的相对路径"
f"(如 'source/xxx.pdf',或用户消息里给的上传路径)。"
)
if resolved.suffix.lower() != ".pdf":
return "", "", (
f"[Error] 仅支持 .pdf(实际 {resolved.suffix or '(无扩展名)'})。"
f"DOCX/PPTX/XLSX 等有文本层的格式请用 markitdown 转换。"
)
try:
raw = resolved.read_bytes()
except OSError as e:
return "", "", f"[Error] 读取文档失败: {type(e).__name__}: {e}"
if len(raw) > max_bytes:
mb = len(raw) / 1024 / 1024
return "", "", (
f"[Error] 文档 {mb:.1f}MB 超过 {max_bytes // 1024 // 1024}MB 上限。"
f"先拆分 / 压缩再传。"
)
b64 = base64.b64encode(raw).decode("ascii")
return f"data:application/pdf;base64,{b64}", display_fn(resolved), ""

275
tools/read_document.py Normal file
View File

@ -0,0 +1,275 @@
"""read_document: 让纯文本主模型""扫描件 / 图片型 PDF。
markitdown 只能抽 PDF 的文本层,扫描件(纯图页)抽出来是空的 这条死路由本 tool
补上:整份 PDF base64 内联喂豆包 seed-2.0-lite 的文档理解(方舟按页栅格化后视觉
OCR),产出 markdown 全文格式 / 上限 / 成本均经 scripts/probe_ark_doc.py 实测:
file 内容块 + `data:application/pdf;base64,` 前缀;单页栅格化 3600 万像素上限;
输入约 1300 token/(100 页约 ¥0.09)
look_at_image 同一 model / key / 记账通道(usage_events kind="vision"),
配置在 config/media/doubao.yaml vision (max_pdf_mb / max_pdf_pages)
"""
from __future__ import annotations
import time
from pathlib import Path
from typing import Any, Optional
from uuid import UUID
from core.ark_client import ArkClient, ArkConfig, ArkError, ArkTimeoutError
from core.storage.usage import record_vision_usage
from .base import Tool, compact_tool_output
from .image_ref import _CONTAINER_ROOT, load_pdf_as_data_url, resolve_in_root
_DEFAULT_QUESTION = (
"这是一份多页 PDF 文档。请逐页把其中的文字完整 OCR 成 markdown:"
"每页以「== 第N页 ==」开头;表格转成 markdown 表格;保留标题层级与段落换行;"
"公式尽量用 LaTeX;不要总结、不要遗漏、不要自行补充原文没有的内容。"
)
# 保存到文件时,tool 返回值里带的正文预览长度(全文在文件里,预览只为让模型确认质量)
_PREVIEW_CHARS = 1500
def _count_pdf_pages(pdf: Path) -> Optional[int]:
"""软探页数(pdfminer 随 markitdown[pdf] 已在依赖里);解析不了返 None 不拦路。"""
try:
from pdfminer.pdfdocument import PDFDocument
from pdfminer.pdfpage import PDFPage
from pdfminer.pdfparser import PDFParser
with open(pdf, "rb") as f:
return sum(1 for _ in PDFPage.create_pages(PDFDocument(PDFParser(f))))
except Exception:
return None
class ReadDocumentTool(Tool):
name = "read_document"
description = (
"Read a PDF (including SCANNED/image-only PDFs that markitdown can't extract) using "
"Doubao Seed 2.0 Lite document understanding — OCRs every page into markdown. "
"Use when markitdown output for a PDF is empty/near-empty (scanned document), or to "
"ask a specific question about a PDF's content. Pass the PDF path; optionally "
"`question` (default: full per-page OCR to markdown) and `save_md` (relative path to "
"write the full text, e.g. 'source/xxx.md' — recommended for multi-page docs so the "
"full text lands in a file instead of flooding context). Costs roughly 0.001-0.002 "
"CNY per page; docs over the page limit must be split first."
)
parameters = {
"type": "object",
"properties": {
"document": {
"type": "string",
"description": (
"PDF 相对路径(task_dir 内,如 'source/xxx.pdf',或用户消息里"
"`[用户上传的文件]` 行给的路径)。"
),
},
"question": {
"type": "string",
"description": (
"想从文档里知道什么(可选)。如「第3章的检测指标是什么」。"
"不传则默认逐页完整 OCR 成 markdown。"
),
},
"save_md": {
"type": "string",
"description": (
"把全文写到这个相对路径(可选,如 'source/xxx.md')。多页 OCR 建议必传:"
"全文落文件,tool 只返回开头预览,不撑爆上下文。"
),
},
},
"required": ["document"],
}
def __init__(
self,
*,
ark_cfg: ArkConfig,
vision_variant_cfg: dict,
variant_key: str,
working_dir: Path,
task_id: UUID,
user_id: UUID,
base_dir: Optional[Path] = None,
user_root: Optional[Path] = None,
) -> None:
super().__init__(base_dir, user_root=user_root)
self.ark_cfg = ark_cfg
self.cfg = vision_variant_cfg
self.variant_key = variant_key
self.working_dir = Path(working_dir)
self.task_id = task_id
self.user_id = user_id
def execute(
self,
document: str,
question: Optional[str] = None,
save_md: Optional[str] = None,
) -> str:
if not (document or "").strip():
return "[Error] document(PDF 路径)不能为空"
cfg = self.cfg
max_bytes = int(float(cfg.get("max_pdf_mb", 30)) * 1024 * 1024)
data_url, disp, err = load_pdf_as_data_url(
document.strip(),
working_dir=self.working_dir,
user_root=self.user_root,
display_fn=self._display,
max_bytes=max_bytes,
)
if err:
return err
# 页数软闸:约 1300 token/页,超过上限会撞模型上下文窗口 → 白付一次失败调用。
# pdfminer 解析不了(加密/损坏)不拦,让 API 报错兜底。
max_pages = int(cfg.get("max_pdf_pages", 100))
resolved = resolve_in_root(document.strip(), self.working_dir, self.user_root)
n_pages = _count_pdf_pages(resolved) if resolved is not None else None
if n_pages is not None and n_pages > max_pages:
return (
f"[Error] 文档 {n_pages} 页超过单次 {max_pages} 页上限(上下文约束)。"
f"先把 PDF 拆成不超过 {max_pages} 页的分卷再逐卷调用。"
)
q = (question or "").strip() or _DEFAULT_QUESTION
model_id = cfg["model_id"]
timeout_s = float(cfg.get("doc_request_timeout_s", 600))
endpoint = cfg.get("endpoint", "/chat/completions")
body: dict[str, Any] = {
"model": model_id,
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": q},
{
"type": "file",
"file": {"filename": Path(disp).name, "file_data": data_url},
},
],
}
],
}
# 超时透明重试,理由同 look_at_image(避免主模型整调用重发、base64 重传)
max_attempts = int(cfg.get("timeout_retries", 1)) + 1
resp = None
for attempt in range(max_attempts):
try:
with ArkClient(self.ark_cfg, timeout_s=timeout_s) as client:
resp = client.post_json(endpoint, body, timeout_s=timeout_s)
break
except ArkTimeoutError as e:
if attempt == max_attempts - 1:
return f"[Error] read_document API: {e}(已重试 {attempt} 次仍超时)"
print(
f"[read_document] timeout, retrying ({attempt + 1}/{max_attempts - 1}): {e}",
flush=True,
)
time.sleep(2 ** attempt)
except ArkError as e:
return f"[Error] read_document API: {e}"
answer, truncated = self._extract_answer(resp)
if not answer:
return (
"[Error] 文档理解响应缺内容(模型未返回文本)。"
"可能 PDF 损坏 / 页面超单页像素上限,稍后重试或先重存该 PDF。"
)
usage = resp.get("usage") or {}
tin = int(usage.get("prompt_tokens", 0) or 0)
tout = int(usage.get("completion_tokens", 0) or 0)
cost_cny = 0.0
try:
cost = record_vision_usage(
task_id=self.task_id,
user_id=self.user_id,
model_profile=f"doubao.{self.variant_key}",
prompt_tokens=tin,
completion_tokens=tout,
input_cny_per_mtoken=float(cfg.get("price_cny_per_mtoken_input", 0)),
output_cny_per_mtoken=float(cfg.get("price_cny_per_mtoken_output", 0)),
extra_units={"document": disp},
)
cost_cny = float(cost)
except Exception as e:
print(f"[read_document] record_vision_usage failed: {type(e).__name__}: {e}", flush=True)
banner = (
f"[read_document] model={model_id} · document={disp}"
+ (f" · pages={n_pages}" if n_pages else "")
+ f" · tokens={tin}+{tout} · cost=¥{cost_cny:.4f}"
)
trunc_note = (
"\n[注意] 输出被模型 token 上限截断,末尾页可能缺失 —— "
"用 question 按页段分次问(如「只 OCR 第 51-100 页」)补齐。"
) if truncated else ""
if save_md and save_md.strip():
saved_disp, save_err = self._save_md(save_md.strip(), answer)
if save_err:
# 保存失败不吞掉 OCR 结果:降级为直接返回(截断保护)
return f"{banner}\n{save_err}(全文改为直接返回)\n\n{compact_tool_output(answer)}{trunc_note}"
preview = answer[:_PREVIEW_CHARS]
more = f"\n...(预览截断,全文 {len(answer)} 字符在文件里)" if len(answer) > _PREVIEW_CHARS else ""
return f"{banner}\nsaved: {saved_disp}\n\n{preview}{more}{trunc_note}"
return f"{banner}\n\n{compact_tool_output(answer)}{trunc_note}"
def _save_md(self, rel: str, text: str) -> tuple[str, str]:
"""把全文写到 task 内。返回 (display_path, error)。
与读取侧 resolve_in_root 同款三形态(相对 / 宿主绝对 / 容器 `/workspace/...`),
但目标是写入 不要求文件已存在,只做 user_root 边界校验
"""
p = Path(rel)
is_container = rel == _CONTAINER_ROOT or rel.startswith(_CONTAINER_ROOT + "/")
if self.user_root is not None and is_container:
target = self.user_root / rel[len(_CONTAINER_ROOT):].lstrip("/")
elif p.is_absolute():
target = p
else:
target = self.working_dir / p
target = target.resolve()
root = (self.user_root or self.working_dir).resolve()
try:
target.relative_to(root)
except ValueError:
return "", f"[Error] save_md 越界: {rel!r}(须在 task 目录内)"
try:
target.parent.mkdir(parents=True, exist_ok=True)
target.write_text(text, encoding="utf-8")
except OSError as e:
return "", f"[Error] 写入 {rel} 失败: {type(e).__name__}: {e}"
return self._display(target), ""
@staticmethod
def _extract_answer(resp: dict) -> tuple[str, bool]:
"""取 choices[0].message.content 文本 + 是否被输出上限截断(finish_reason=length)。"""
choices = resp.get("choices")
if not (isinstance(choices, list) and choices and isinstance(choices[0], dict)):
return "", False
truncated = choices[0].get("finish_reason") == "length"
msg = choices[0].get("message")
if not isinstance(msg, dict):
return "", truncated
content = msg.get("content")
if isinstance(content, str):
return content.strip(), truncated
if isinstance(content, list):
parts = [
c.get("text", "")
for c in content
if isinstance(c, dict) and c.get("type") == "text"
]
return "\n".join(p for p in parts if p).strip(), truncated
return "", truncated

View File

@ -691,7 +691,7 @@ async function loadEarlierMessages() {
if (earlier.length) state.loadedMessages = earlier.concat(state.loadedMessages);
state.msgHasMore = !!data.has_more;
state.msgLoadingEarlier = false;
renderMessages(state.loadedMessages);
renderMessages(state.loadedMessages, { stickBottom: false });
// 锚回:新增内容都在上方,保持原先可见的首条仍在原位
wrap.scrollTop = prevTop + (wrap.scrollHeight - prevH);
} catch (e) {
@ -720,7 +720,7 @@ async function loadNewerMessages() {
if (newer.length) state.loadedMessages = state.loadedMessages.concat(newer);
state.msgHasMoreNewer = !!data.has_more_after;
state.msgLoadingNewer = false;
renderMessages(state.loadedMessages);
renderMessages(state.loadedMessages, { stickBottom: false });
wrap.scrollTop = prevTop; // 新增在下方,保持原视口不跳
} catch (e) {
state.msgLoadingNewer = false;
@ -755,19 +755,23 @@ async function loadMessagesAround(idx) {
state.msgHasMoreNewer = !!data.has_more_after;
state.msgLoadingEarlier = false;
state.msgLoadingNewer = false;
renderMessages(state.loadedMessages);
// 不滚底:jumpToMessage 随后瞬时定位到目标卡;滚底会把视口钉到窗口末尾,
// 底部 sentinel 立刻进视口触发补载重渲染,打断跳转(首次点圆点跳不到位的根因)
renderMessages(state.loadedMessages, { stickBottom: false });
return true;
}
async function jumpToMessage(idx) {
const wrap = $("chat-stream");
let card = wrap.querySelector(`.msg[data-idx="${idx}"]`);
let rebuilt = false; // 本次跳转是否重建了消息窗口
if (!card) {
let ok = false;
try { ok = await loadMessagesAround(idx); }
catch (e) { if (e.status === 401) { logout(); return; } }
if (!ok) return;
card = wrap.querySelector(`.msg[data-idx="${idx}"]`);
rebuilt = true;
}
if (!card) return;
// 顶部对齐(非居中):第一轮上方无内容无法居中、会被钉到顶端,而 updateActiveOutlineDot
@ -775,7 +779,10 @@ async function jumpToMessage(idx) {
// .msg 的 scroll-margin-top 给卡片留一点上方呼吸空间。
setActiveOutlineIdx(idx);
lockOutlineDuringJump(); // 锁住活跃圆点:平滑滚动途中的 scroll 事件不得把活跃态抢到途经轮次
card.scrollIntoView({ behavior: "smooth", block: "start" });
// 重建窗口后瞬时定位:视口位置还停在旧窗口的残留值,平滑动画要横穿整窗,
// 途中 sentinel 补载 / 图片异步撑高都会让落点偏掉;瞬时跳没有窗口期。
// 已在 DOM 的近距离跳转保留平滑。
card.scrollIntoView({ behavior: rebuilt ? "auto" : "smooth", block: "start" });
card.classList.add("msg-jump-flash");
setTimeout(() => card.classList.remove("msg-jump-flash"), 1200);
}
@ -783,6 +790,10 @@ async function jumpToMessage(idx) {
// 顶/底 sentinel 进视口即自动补更早 / 更新 —— 复用 task list 的同款范式。
// root 是 chat-stream 滚动容器;每次 renderMessages 重建 DOM 后重新 observe 新 sentinel。
const _msgScrollObserver = new IntersectionObserver((entries) => {
// 目录跳转动画期间不自动补载:补载的整窗重渲染会删掉平滑滚动的目标元素、
// 把视口拽回 fetch 前的位置(表现:点圆点跳不到位)。解锁时 _repokeMsgSentinels
// 会补投一次当前交叉状态,不会漏加载。
if (_outlineJumpLock) return;
for (const en of entries) {
if (!en.isIntersecting) continue;
if (en.target.classList.contains("msg-top-sentinel")) {
@ -793,6 +804,16 @@ const _msgScrollObserver = new IntersectionObserver((entries) => {
}
}, { root: $("chat-stream"), rootMargin: "150px 0px" });
// IntersectionObserver 只在交叉状态变化时通知 —— 锁定期吞掉的通知不会重发,
// 若 sentinel 一直停在视口内,解锁后就没人触发补载了。unobserve+observe
// 强制 observer 重新投递一次当前交叉状态,补上这一拍。
function _repokeMsgSentinels() {
$("chat-stream").querySelectorAll(".msg-top-sentinel, .msg-bot-sentinel").forEach((el) => {
_msgScrollObserver.unobserve(el);
_msgScrollObserver.observe(el);
});
}
// ───── 消息目录(右侧悬浮圆点轨道)─────
// 切 task / run 收尾后拉全部 user 轮次;点圆点 jumpToMessage 定位;滚动时高亮当前轮。
async function refreshOutline() {
@ -847,6 +868,7 @@ function lockOutlineDuringJump() {
_outlineJumpTimer = setTimeout(() => {
_outlineJumpLock = false;
updateActiveOutlineDot();
_repokeMsgSentinels(); // 锁定期吞掉的 sentinel 通知补投一次(见 observer 注释)
}, 700);
}
@ -1149,7 +1171,10 @@ async function copyTextToClipboard(text, btn) {
}
}
function renderMessages(msgs) {
// stickBottom:渲染完是否滚到底。首屏/发送等"看最新"路径保持默认 true;
// 调窗口的路径(向上/向下补载、目录跳转)传 false —— 它们各自负责定位视口,
// 中间插一次滚底会触发多余 scroll 事件,跳转路径还会把视口钉错位置(0.58.55 修)。
function renderMessages(msgs, { stickBottom = true } = {}) {
const wrap = $("chat-stream");
_msgScrollObserver.disconnect(); // 旧 sentinel 随 innerHTML 清掉,先断开避免悬挂 observe
wrap.innerHTML = "";
@ -1302,7 +1327,7 @@ function renderMessages(msgs) {
wrap.appendChild(sb);
_msgScrollObserver.observe(sb);
}
wrap.scrollTop = wrap.scrollHeight;
if (stickBottom) wrap.scrollTop = wrap.scrollHeight;
setTaskProgress(state.taskId, currentProgressSteps);
upgradeMediaArtifacts(wrap);
renderPersistedRunTerminal(); // 上次 run error/cancelled 终态 → 末尾补持久卡(所有重渲路径统一走这)