适合边界清晰的短任务Best for bounded tasks
- 上下文集中、成本和延迟较低Central context, lower cost and latency
- 固定字段查询、公式计算、单文档提取Known fields, formulas, single-file extraction
- 错误来源和权限较难拆分Harder to isolate errors and privileges
7个同步调查Agent+1个异步学习档案Agent+1个人工审批节点。系统只组织证据、生成待验证假设并管理风险,不替代工程师确认物理根因。Seven synchronous investigation agents, one asynchronous learning agent and one human approval gate. The system organizes evidence and testable hypotheses; it never replaces engineering causality decisions.
理论作业一THEORY 01
单Agent在一个上下文内完成理解、规划、工具调用和输出;多Agent把复杂目标拆为带独立职责、最小权限、输入输出契约和审查关系的任务单元。只有单Agent基线暴露出跨专业、权限隔离或独立复核问题时,复杂度才有理由增加。A single agent works end-to-end in one context. A multi-agent system splits the goal into bounded tasks with independent roles, least privilege, contracts and review. Added complexity is justified only when a single-agent baseline fails on specialization, permission isolation or independent verification.
理论作业二THEORY 02
状态机决定下一步和停止条件;Agent只在明确任务内生成候选。风险审查失败、来源不足、接口超时或权限不符时必须HOLD并转责任工程师。The state machine controls transitions and stop rules; agents propose only within assigned tasks. Missing evidence, risk failure, timeout or denied access forces HOLD and human handoff.
研究员与报告分析员顺序交接,并包装成对外服务。Researcher-to-analyst sequential handoff exposed as a service.
展示角色分工、知识检索、结构化生成的组合。Role specialization, retrieval and structured generation.
编码案例仍出现运行报错,因此必须加入确定性测试、失败回放和人工门槛。A generated program still failed at runtime, requiring deterministic tests, replay and human gates.
现场落地核验FIELD DEPLOYMENT READINESS
下列平台均是拟接入或核验对象,不冒充TCL华星现状。真实协作入口、接口、字段字典、时钟同步、标签口径与责任人,必须通过访谈和授权确认。The following systems are proposed verification targets—not claims about TCL CSOT's current stack. Collaboration entry, interfaces, field dictionaries, clock alignment, label definitions and owners require interviews and authorization.
实践作业一PRACTICE 01
| 交接Handoff | 量化门槛Gate | 失败处理Failure action | 证据状态Evidence |
|---|---|---|---|
| Ticket → A1 | 关键字段完整率≥95%;缺失显式率100%Key-field completeness ≥95%; missingness explicit 100% | HOLD · 补问,不猜测ask, never guess | PILOT TARGET |
| A1 → A2/A4 | 范围字段100%;追溯键与时区明确100%Scope 100%; trace key and timezone explicit 100% | 退回范围确认Return scope | PILOT TARGET |
| A2 → A3 | 血缘100%;时钟偏差在已批准容差内;缺失显式100%Provenance 100%; clock drift within approved tolerance; missingness explicit 100% | DATA_MISSING / HOLD | PILOT TARGET |
| A3 + A4 → A5 | 复算100%;文档ID/版本/锚点100%;语义映射经Owner确认Recompute 100%; doc ID/version/anchor 100%; semantic mapping owner-approved | REJECT / STALE_KNOWLEDGE | PILOT TARGET |
| A5 → A6 | ≥1支持+≥1反证/未知;无依据因果=0≥1 support + ≥1 challenge/unknown; unsupported causality=0 | REJECT | PILOT TARGET |
| A6 → A7 | 原始追溯/重跑参数/异议入口100%Raw trace/rerun parameters/dispute entry 100% | PACKAGE_INCOMPLETE | PILOT TARGET |
| A7 → H1 | 引用≥98%;危险动作0;L1—L4门槛100%Citation ≥98%; dangerous actions 0; L1–L4 gate 100% | HOLD / REJECT | PILOT TARGET |
| H1 → A8 | Owner/结论/证据/原因/版本/时间100%;QA另行批准Owner/decision/evidence/reason/version/time 100%; separate QA approval | memory_write=PENDING_QA | PILOT TARGET |
真实公开案例回放REAL PUBLIC CASE REPLAY
UCI SECOM原始失败样本,仅验证分析管线、来源记录和停止边界;不证明TCL效果,也不证明物理根因。A real UCI SECOM failure used only to verify pipeline, provenance and stop behavior—not TCL impact or physical causality.
当前只能证明管线在不同失败模式下能检测、降级和停止,不能证明对真实工厂物理根因的泛化。真实效果评估必须取得授权历史案例,按失效族、设备族、班次和工艺版本分层,并按时间或设备留出测试集,防止同源泄漏。This proves detection, degradation and stopping across failure modes—not generalization to physical factory causality. Real evaluation requires authorized cases stratified by failure family, equipment, shift and process version, with time/equipment holdout to prevent leakage.
实践作业二PRACTICE 02
“结果离谱、界面复杂、响应慢”是题设反馈。真实上线前应先冻结Before基线;当前工厂数值、接口准备度与After效果均待企业授权采集。“Absurd results, complex UI and slow responses” are assignment scenarios. A real launch first freezes a Before baseline; factory values, interface readiness and After impact remain pending authorization.
8/10目标用户在无培训条件下完成范围确认、证据查看与人工裁决;核心任务无横向滚动、无隐藏关键风险、移动端可操作。8/10 target users complete scope confirmation, evidence review and decision without training; no hidden risk, horizontal scroll or unusable mobile action.
思考作业一REFLECTION 01
闭环为:失败案例→人工裁决→质量审核→冻结评测集→版本回归。优先级先看安全和权限,再看核心任务质量、用户负担,最后才是美观与新模型。The loop is failure → human decision → QA → frozen evaluation set → version regression. Prioritize safety and access, then core quality and user burden, and only then polish or new models.
思考作业二REFLECTION 02
跨记录检索、统一口径、比较窗口、发现重复模式、保存来源与差异。AI输出是候选证据,不是物理因果。Search records, normalize terms, compare windows, detect repeats and preserve evidence. Outputs are candidates, not physical causality.
补充设备声音、维护切换、材料变化、临时工艺等系统外信息;接受、暂缓或覆盖AI建议,并签署理由。Add off-system context such as sound, maintenance, material and process changes; accept, hold or override with signed rationale.
思考作业三 · 交付REFLECTION 03 · DELIVERY
把MES、SCADA、设备日志、检测、SOP和沟通中的分散信息,对齐为带来源、时间、版本和责任人的证据包,让工程师把时间用于判断,而不是找资料。Align dispersed MES, SCADA, log, inspection, SOP and communication data into a traceable package so engineers spend time judging—not searching.
多Agent、自动化与知识沉淀可能让同源偏差形成错误共识,并以更快速度、更高权威感扩散。控制手段是来源、反证、人工门槛、只读权限和审核后学习。Multi-agent workflows can turn correlated bias into false consensus and spread it faster. Controls are provenance, challenge evidence, human gates, read-only access and QA-approved learning.
公式同时扣除Agent准备、人工复核、模型系统与误报处理成本。所有默认值仅用于展示测量结构。The formula deducts agent preparation, human review, system/model and false-alert costs. Defaults only demonstrate the measurement structure.
| 变量Variable | 操作定义Operational definition | 首选数据源Preferred source | 校准责任人Owner | 当前状态Status |
|---|---|---|---|---|
| N | 按统一口径进入调查的月案例数,去重后计数Deduplicated monthly investigation cases under one definition | MES / QMS / incident ticket | 业务OwnerBusiness owner | NOT MEASURED |
| T_before | 从责任人接单到首份可复核证据清单Assignment to first reviewable evidence list | 工单时间戳+两周时间研究Ticket timestamps + two-week time study | 产线工程师Line engineer | PENDING BASELINE |
| T_agent / T_review | 系统执行时长与人工复核时长分开记录System run and human review measured separately | 不可变运行轨迹+复核事件Immutable traces + review events | 产品+质量Product + Quality | SHADOW REQUIRED |
| C_hour | 含福利与间接成本的财务批准口径Finance-approved loaded labor rate | 财务成本口径Finance cost standard | 财务Finance | PENDING FINANCE |
| C_model_infra | 模型、算力、存储、运维与审计留存的全成本Model, compute, storage, ops and audit retention | 账单+容量预算Billing + capacity plan | IT / FinOps | SCENARIO DEFAULT |
| F / T_false | 影子期误报数及其实际处理时长Shadow false alerts and actual handling time | 盲评标签+处理事件Blind labels + handling events | 质量OwnerQuality owner | SHADOW REQUIRED |
| 停机/质量收益Downtime/quality benefit | 需有同期对照或可信反事实;当前公式明确排除Requires control/counterfactual; explicitly excluded now | MES / OEE / QMS | 生产+财务Production + Finance | EXCLUDED |