第19天 · 多Agent产品系统DAY 19 · MULTI-AGENT PRODUCT SYSTEM

从单体判断到协同调查FROM SIGNAL TO SHARED EVIDENCE制造良率调查多Agent系统MANUFACTURING YIELD MULTI-AGENT SYSTEM

7个同步调查Agent+1个异步学习档案Agent+1个人工审批节点。系统只组织证据、生成待验证假设并管理风险,不替代工程师确认物理根因。Seven synchronous investigation agents, one asynchronous learning agent and one human approval gate. The system organizes evidence and testable hypotheses; it never replaces engineering causality decisions.

A1—A7 · SYNCA8 · ASYNCH1 · ACCOUNTABLE确定性流程+有限推理Deterministic flow + bounded reasoning
证据状态严格分离:课程案例、公开数据复算、产品试点目标、待企业授权采集。页面中的目标值均不冒充TCL实绩。Evidence states are separated: course cases, recomputed public data, pilot targets and enterprise data pending authorization. Targets are never presented as TCL results.
靠近节点自动吸附 · 点击跳转对应章节Move near a node to magnetize · click to jump
COURSE EVIDENCE20—23集案例与机制Episodes 20–23
PUBLIC RECOMPUTEDUCI SECOM失败样本UCI SECOM failure
PILOT TARGET门槛,不是已达成绩Gates, not results
PENDING AUTHORIZATION企业接口、基线与ROIInterfaces, baseline, ROI
01

理论作业一THEORY 01

多Agent的差异不在数量,而在可验证的责任拆分Multi-agent value comes from verifiable responsibility—not agent count

单Agent在一个上下文内完成理解、规划、工具调用和输出;多Agent把复杂目标拆为带独立职责、最小权限、输入输出契约和审查关系的任务单元。只有单Agent基线暴露出跨专业、权限隔离或独立复核问题时,复杂度才有理由增加。A single agent works end-to-end in one context. A multi-agent system splits the goal into bounded tasks with independent roles, least privilege, contracts and review. Added complexity is justified only when a single-agent baseline fails on specialization, permission isolation or independent verification.

SINGLE AGENT

适合边界清晰的短任务Best for bounded tasks

  • 上下文集中、成本和延迟较低Central context, lower cost and latency
  • 固定字段查询、公式计算、单文档提取Known fields, formulas, single-file extraction
  • 错误来源和权限较难拆分Harder to isolate errors and privileges
MULTI AGENT

适合跨专业、高风险、长流程Best for cross-domain, high-risk work

  • 职责、上下文和工具权限可隔离Roles, context and tool access are isolated
  • 可并行、可交接、可独立反证Parallel, contract-based and independently challenged
  • 增加编排、调用、测试和失败处理成本Adds orchestration, test and failure costs
01 · 专业差异Specialization数据、知识与风险审查能力不同Data, knowledge and risk skills differ
02 · 权限隔离Permission isolationMES只读与SOP检索必须分权MES read access differs from SOP retrieval
03 · 独立反证Independent challenge生成者不能自我批准高风险结论The author cannot approve its own claim
04 · 交接可验收Testable handoffs每一步能输出结构化契约Every step has a structured contract
05 · 可安全并行Safe parallelism数据检查与SOP检索互不依赖Evidence and SOP search can run in parallel
06 · 审计责任Auditability错误能定位到Agent、工具与版本Errors trace to agent, tool and version
采用规则Adoption rule至少满足三项条件,且单Agent盲测基线已经暴露相应问题,才进入多Agent方案。多个相似模型重复回答同一问题不构成质量控制。Adopt multi-agent only when at least three criteria apply and a single-agent blind baseline shows the corresponding failure. Repeating the same question across similar models is not quality control.
02

理论作业二THEORY 02

CrewAI:像组建专业团队一样定义角色、任务、工具和流程CrewAI defines a professional team through roles, tasks, tools and process

AGENT角色·目标·权限Role · goal · access
TASK工作·产物·验收Work · output · acceptance
TOOL外部能力External capability
PROCESS顺序·层级·状态Order · hierarchy · state
CREW一次调查团队One investigation team
FASTAPI服务入口Service boundary
本项目不是自由自治讨论,而是“确定性外层流程+有限Agent推理”This system uses a deterministic outer process with bounded agent reasoning

状态机决定下一步和停止条件;Agent只在明确任务内生成候选。风险审查失败、来源不足、接口超时或权限不符时必须HOLD并转责任工程师。The state machine controls transitions and stop rules; agents propose only within assigned tasks. Missing evidence, risk failure, timeout or denied access forces HOLD and human handoff.

COURSE 20

CrewAI + FastAPI

研究员与报告分析员顺序交接,并包装成对外服务。Researcher-to-analyst sequential handoff exposed as a service.

COURSE 21–22

研究与RAG案例Research and RAG cases

展示角色分工、知识检索、结构化生成的组合。Role specialization, retrieval and structured generation.

COURSE 23

多Agent不等于自动正确More agents do not guarantee correctness

编码案例仍出现运行报错,因此必须加入确定性测试、失败回放和人工门槛。A generated program still failed at runtime, requiring deterministic tests, replay and human gates.

03

现场落地核验FIELD DEPLOYMENT READINESS

先确认TCL岗位、平台和门槛,再谈Agent进入现场Verify TCL roles, platforms and gates before introducing agents

下列平台均是拟接入或核验对象,不冒充TCL华星现状。真实协作入口、接口、字段字典、时钟同步、标签口径与责任人,必须通过访谈和授权确认。The following systems are proposed verification targets—not claims about TCL CSOT's current stack. Collaboration entry, interfaces, field dictionaries, clock alignment, label definitions and owners require interviews and authorization.

PENDING VERIFICATION

六道落地门槛 · 当前先补G1与G2Six deployment gates · address G1 and G2 first

G1 · P0试点范围/Owner
部分完成
Pilot scope/owner
Partial
G2 · P0平台/数据接入
未满足
Platform/data
Not met
G3 · P1Before基线
未测量
Before baseline
Unmeasured
G4 · P2授权离线评估
仅公开管线
Authorized offline
Public only
G5 · P3影子运行
未开始
Shadow mode
Not started
G6 · P4工程师辅助
未授权
Engineer assist
Unauthorized
04

实践作业一PRACTICE 01

七个同步Agent+一个异步学习Agent,H1保留唯一业务裁决权Seven synchronous agents plus one asynchronous learner; H1 owns the decision

A1
SYNC

显式协作拓扑:采集、分析、解释、编排与审查互不替代Explicit topology: collection, analysis, explanation, packaging and review stay distinct

A1 · ORCHESTRATE范围与任务编排Scope and task orchestration校验批次、工序、时间窗;不判断根因。Validate lot, step and window; no root-cause claim.
A2 · COLLECT数据采集与血缘Collection and lineage只读拉取、保留原始键、时钟与版本。Read-only retrieval with keys, clocks and versions.
A4 · DOMAINSOP与工艺语义SOP and process semantics只检索有效授权知识,失效版本不得引用。Retrieve only valid authorized knowledge.
→STRUCTURED
HANDOFF
A3 · ANALYZE统计分析Statistical analysis复算、比较窗口、异常排序;统计关联不是因果。Recompute, compare and rank; association is not causality.
A5 · FALSIFY假设与反证Hypothesis and falsification最多三条候选;每条必须有支持、反证或未知。At most three candidates, each with support and challenge/unknown.
A6 · PACKAGE证据编排与报告Evidence packaging and report生成可追溯、可重跑、可标异议的调查包。Build a traceable, rerunnable and disputable package.
→INDEPENDENT
REVIEW
A7 · RISK权限、来源与措辞审查Access, provenance and claim review可PASS/HOLD/REJECT;不可执行生产动作。May PASS/HOLD/REJECT; cannot act on production.
H1 · ACCOUNTABLE HUMAN责任工程师裁决Accountable engineer decision接受、暂缓或覆盖并签署理由;AI无裁决权。Accept, hold or override with rationale; AI has no decision authority.
A8 · ASYNC MEMORY结案后候选学习Post-closure candidate learning仅QA批准后进入版本化知识;不参与当前案件判断。Versioned knowledge only after QA; never influences the active case.
复杂度门槛:试点必须用同案盲评证明拆分带来更完整证据、更低编辑负担或更强权限隔离;若A2+A3或A6+A7合并后无显著退化,就合并角色,避免为“多Agent”而多Agent。Complexity gate: blind same-case evaluation must show better evidence, lower edit burden or stronger access isolation. Merge A2+A3 or A6+A7 when separation adds no measurable value.

协作状态机 · 点击运行到人工门槛Collaboration state machine · run until the human gate

CASE-SECOM-F-0015

PUBLIC REPLAY

异常不是报错页,而是有责任人、有退出条件的业务分支Exceptions are owned business branches with exit conditions—not error pages

INVALID_INPUT

Agent交接量化门槛 · 全部为试点目标Quantified handoff gates · all are pilot targets

交接Handoff量化门槛Gate失败处理Failure action证据状态Evidence
Ticket → A1关键字段完整率≥95%;缺失显式率100%Key-field completeness ≥95%; missingness explicit 100%HOLD · 补问,不猜测ask, never guessPILOT TARGET
A1 → A2/A4范围字段100%;追溯键与时区明确100%Scope 100%; trace key and timezone explicit 100%退回范围确认Return scopePILOT TARGET
A2 → A3血缘100%;时钟偏差在已批准容差内;缺失显式100%Provenance 100%; clock drift within approved tolerance; missingness explicit 100%DATA_MISSING / HOLDPILOT TARGET
A3 + A4 → A5复算100%;文档ID/版本/锚点100%;语义映射经Owner确认Recompute 100%; doc ID/version/anchor 100%; semantic mapping owner-approvedREJECT / STALE_KNOWLEDGEPILOT TARGET
A5 → A6≥1支持+≥1反证/未知;无依据因果=0≥1 support + ≥1 challenge/unknown; unsupported causality=0REJECTPILOT TARGET
A6 → A7原始追溯/重跑参数/异议入口100%Raw trace/rerun parameters/dispute entry 100%PACKAGE_INCOMPLETEPILOT TARGET
A7 → H1引用≥98%;危险动作0;L1—L4门槛100%Citation ≥98%; dangerous actions 0; L1–L4 gate 100%HOLD / REJECTPILOT TARGET
H1 → A8Owner/结论/证据/原因/版本/时间100%;QA另行批准Owner/decision/evidence/reason/version/time 100%; separate QA approvalmemory_write=PENDING_QAPILOT TARGET
05

真实公开案例回放REAL PUBLIC CASE REPLAY

证据不足时安全停止,比“硬给答案”更接近真实落地Safe stopping under insufficient evidence is more deployable than forced answers

PUBLIC RECOMPUTED

CASE-SECOM-F-0015

UCI SECOM原始失败样本,仅验证分析管线、来源记录和停止边界;不证明TCL效果,也不证明物理根因。A real UCI SECOM failure used only to verify pipeline, provenance and stop behavior—not TCL impact or physical causality.

15SOURCE ROW
FAILLABEL = 1
28 / 590MISSING FIELDS
+11.90FEATURE_431 Z
回放结论Replay conclusionA3只报告可复算的统计关联,A4因缺少设备语义而拒绝编造映射,A5仅保留候选核查项,A6生成可追溯证据包,A7将root_cause保持null并HOLD。该案例证明边界与停止机制,不证明多Agent优于单Agent。A3 reports only reproducible association, A4 refuses invented semantics, A5 keeps a candidate check, A6 packages traceable evidence and A7 holds with root_cause=null. This proves boundaries and stopping—not superiority over a single agent.

五种失败模式验证:一例真实公开回放+四例受控故障注入Five failure modes: one real public replay plus four controlled fault injections

当前只能证明管线在不同失败模式下能检测、降级和停止,不能证明对真实工厂物理根因的泛化。真实效果评估必须取得授权历史案例,按失效族、设备族、班次和工艺版本分层,并按时间或设备留出测试集,防止同源泄漏。This proves detection, degradation and stopping across failure modes—not generalization to physical factory causality. Real evaluation requires authorized cases stratified by failure family, equipment, shift and process version, with time/equipment holdout to prevent leakage.

泛化验收的候选采样下限(待业务Owner校准):Candidate sampling floor for generalization, pending owner calibration: 不少于30个授权历史案件,并保证每个纳入评估的主要失效族至少10例;训练/检索与测试按时间或设备隔离。若样本不足,只报告个案结果与置信区间,不宣称泛化。At least 30 authorized historical cases and 10 per included major failure family; isolate training/retrieval and test by time or equipment. With fewer cases, report case-level outcomes and uncertainty only.

Top-5不是“五个看起来显著的数”,而是五条待证伪的工艺假设契约Top-5 is not five impressive numbers; it is five falsifiable process-hypothesis contracts

RANK 1 · PUBLICfeature_431 · z=+11.90UCI字段匿名,无法诚实映射工艺。下一步:追溯字段字典、传感器/计算口径、设备谱系;未获得语义前root_cause=null。UCI field is anonymous. Trace dictionary, sensor/formula and equipment lineage; root_cause remains null.
RANK 2 · PENDING待授权数据计算Pending authorized compute必须同时记录效应量、时间窗稳定性、缺失模式、可能工艺机制和反证。Record effect size, temporal stability, missingness, mechanism and counterevidence.
RANK 3 · PENDING待授权数据计算Pending authorized compute“统计显著”不得直接转换成“设备故障”,需工艺Owner确认语义。Statistical significance cannot become equipment failure without domain-owner semantics.
RANK 4 · PENDING待授权数据计算Pending authorized compute需检查是否为共同上游、批次结构、维护切换或测量漂移造成的伪相关。Test confounding from upstream, lot structure, maintenance or measurement drift.
RANK 5 · PENDING待授权数据计算Pending authorized compute只有现场验证动作成功、反例未推翻且工程师签字,才可提升为可复用结论。Promotion requires a successful field test, surviving counterexamples and engineer sign-off.
工艺解释模板(不是UCI字段映射):压力信号→腔体/阀门假设;温度信号→加热区假设;流量信号→MFC/供气假设;功率信号→匹配网络假设;量测信号→沉积均匀性/量测漂移假设。每条都必须配套“如何证伪”。Domain template—not UCI mapping: pressure→chamber/valve; temperature→heater zone; flow→MFC/supply; power→matching network; metrology→uniformity/drift. Every hypothesis needs a falsification test.

证据包的操作定义:能追溯、能重跑、能比较、能标异议Operational evidence package: trace, rerun, compare and dispute

当前为前端原型交互:展示操作语义与审计字段,不声称已连接企业后端。生产实现必须使用授权身份、不可变事件日志、数据版本与分析版本;“重跑”生成新版本,绝不覆盖原输出。This is a front-end interaction prototype. Production requires authorized identity, immutable events, data and analysis versions. Reruns create new versions and never overwrite originals.
06

实践作业二PRACTICE 02

上线一个月反馈不是结果,而是诊断入口Month-one feedback is a diagnostic input—not a result

“结果离谱、界面复杂、响应慢”是题设反馈。真实上线前应先冻结Before基线;当前工厂数值、接口准备度与After效果均待企业授权采集。“Absurd results, complex UI and slow responses” are assignment scenarios. A real launch first freezes a Before baseline; factory values, interface readiness and After impact remain pending authorization.

P0 · SAFETY

先冻结动作,再分型定位Freeze actions, then classify failures

  1. 保持只读并保存Agent轨迹、输入、工具结果和模型版本Stay read-only and preserve traces, inputs, tool outputs and versions
  2. 按范围/时间窗、数据错位、引用、推理、知识版本、反馈口径六类标注Label scope, alignment, citation, reasoning, knowledge and feedback failures
  3. 失败案例经质量审核进入冻结评测集并做版本回归QA-approved failures enter a frozen evaluation set and regression suite
PILOT TARGET

恢复门槛Recovery gates

≥95%失败召回Failure recall
≥98%引用准确率Citation accuracy
100%来源覆盖Provenance coverage
0危险动作Dangerous actions
P1 · USABILITY

按决策顺序收敛默认界面Reduce the default view to the decision sequence

  1. 默认只显示异常摘要、AI候选、关键证据、工程师操作Show summary, candidates, key evidence and human actions only
  2. Agent轨迹、完整字段和日志渐进展开Progressively disclose traces, fields and logs
  3. 核心操作统一为接受、驳回、查看证据,并保留键盘可达性Standardize accept, reject and view-evidence actions with keyboard access
PILOT TARGET

可用性验收Usability acceptance

8/10目标用户在无培训条件下完成范围确认、证据查看与人工裁决;核心任务无横向滚动、无隐藏关键风险、移动端可操作。8/10 target users complete scope confirmation, evidence review and decision without training; no hidden risk, horizontal scroll or unusable mobile action.

P1 · PERFORMANCE

先做分节点P50/P95,再优化架构Measure per-node P50/P95 before redesign

  1. A2与A3并行;确定性规则和缓存优先于LLMParallelize A2/A3; prefer rules and cache before LLMs
  2. 压缩交接Schema、按难度路由模型、限制重试和上下文Compress contracts, route models by difficulty and bound retries/context
  3. 1秒内返回状态,计划和首批证据先行,完整证据包异步通知Return state within one second; stream plan and first evidence; notify when complete
PILOT TARGET

性能预算Latency budget

≤1sUI FEEDBACK
≤10sPLAN P95
≤30sFIRST EVIDENCE P95
≤120sFULL PACKAGE P95

上线前Before基线与数据准备度Before baseline and data readiness before launch

至少2周或30例At least 2 weeks or 30 cases异常量、准备时间P50/P90、首报、缺失率、重复沟通、闭环周期Volume, prep P50/P90, first report, missingness, contacts, closure time
MES / SCADA / LOG / QMS接口、格式、权限、时钟、追溯键、责任人均待确认Interfaces, format, access, clocks, trace keys and owners pending
当前状态Current state未测量 / 待企业授权采集;不得填入目标值冒充基线Unmeasured / pending authorization; never use targets as baselines

单Agent vs 多Agent:同案、盲评、同量表Single vs multi: same cases, blind experts, one rubric

ARM A · DAY18SINGLE

  • 同一批授权历史案例Same authorized historical cases
  • 同一数据权限和时间预算Same access and time budget
  • 责任工程师不知道输出来源Engineers are blinded to system identity

ARM B · DAY19MULTI

  • 证据完整≥95%,来源100%Evidence ≥95%, provenance 100%
  • 盲评通过率较A组+10个百分点Blind pass rate +10pp over A
  • 编辑率≤20%或较A下降20%,P95≤120秒Edit rate ≤20% or -20% vs A; P95 ≤120s
GO / NO-GO: 无依据因果=0、高风险动作=0、净增量价值为正;任何危险建议一次即No-Go。当前尚未获得授权案例,未得出优胜结论。Unsupported causality=0, risky actions=0 and positive incremental value. One dangerous suggestion triggers No-Go. No winner is claimed before authorized evaluation.
07

思考作业一REFLECTION 01

最关键的持续优化,是把失败变成可回归验证的资产The most important optimization turns failures into regression assets

闭环为:失败案例→人工裁决→质量审核→冻结评测集→版本回归。优先级先看安全和权限,再看核心任务质量、用户负担,最后才是美观与新模型。The loop is failure → human decision → QA → frozen evaluation set → version regression. Prioritize safety and access, then core quality and user burden, and only then polish or new models.

P0安全与越权Safety & access错误因果、危险动作、审计缺口Causal overclaim, risky action, audit gap
P1核心质量Core quality漏报、引用、证据完整、响应Misses, citations, evidence, latency
P2用户负担User burden编辑、重复沟通、可用性Edits, contacts, usability
P3增强体验Enhancement文案、动效、低频功能、新模型Copy, motion, low-frequency features
没有进入真实流程的先进性无法产生价值;没有可测净收益的落地无法持续。因此决策顺序必须是:可落地性 → 可验证ROI → 技术先进性。Technology outside a real workflow creates no value; deployment without measurable net benefit cannot persist. The decision order is deployability → verifiable ROI → technical sophistication.先进性是加速器,不是上线许可证。Sophistication is an accelerator—not a license to deploy.

L1—L4动作风险门槛:级别越高,Agent权限越低L1–L4 action-risk gates: higher risk means less agent authority

L1信息与检索Information查SOP、列证据、说明缺口Retrieve SOP, list evidence and gaps允许:只读+引用;门槛:来源覆盖100%。Allowed read-only with 100% provenance.
L2调查建议Investigation advice建议核查窗口、信号、维护记录Suggest checks of windows, signals and maintenance允许:候选核查清单;必须人工确认,root_cause=null。Candidate checklist only; human review; root_cause=null.
L3影响生产Production impact改配方、停线、放行、判废Recipe, stop, release or scrap本产品阻断;即使人工提出也需脱离Agent、走正式审批与安全评审。Blocked; requires formal off-agent approval and safety review.
L4安全与合规控制Safety and compliance联锁、EHS、法定检验Interlocks, EHS and statutory inspection永久禁止:Agent不得建议绕过、执行或弱化。Permanently prohibited: no bypass, execution or weakening.

四阶段先验证再放权Four-stage validate-before-authorize path

永久禁止:设备/配方写入、自动停线、自动放行/判废、绕过安全联锁。Permanently prohibited: equipment/recipe writes, automatic line stop, release/scrap and safety-interlock bypass.
08

思考作业二REFLECTION 02

日常调查AI辅助人,高风险判断人监督AIAI assists routine investigation; people supervise high-risk judgment

AI

扩大观察范围Expand observation

跨记录检索、统一口径、比较窗口、发现重复模式、保存来源与差异。AI输出是候选证据,不是物理因果。Search records, normalize terms, compare windows, detect repeats and preserve evidence. Outputs are candidates, not physical causality.

ENGINEER

承担因果与风险责任Own causality and risk

补充设备声音、维护切换、材料变化、临时工艺等系统外信息;接受、暂缓或覆盖AI建议,并签署理由。Add off-system context such as sound, maintenance, material and process changes; accept, hold or override with signed rationale.

AI与工程师冲突演练AI–engineer conflict drill

FREEZE冻结动作Freeze action
PRESERVE保留原输出Preserve output
COMPARE并排证据Compare evidence
DECIDE责任人裁决Owner decides
VERIFY质量复核QA verifies
LEARN审核后学习Learn after QA
status=HOLD · action=READ_ONLY · ai_output=IMMUTABLE · memory_write=PENDING_QA
A8 · ASYNC · PENDING_QA

老师傅经验必须先变成可核验的学习档案Expert experience must become a verifiable learning record

SYMPTOM症状与工序Symptom and step
PRECONDITION适用设备与前提Equipment and conditions
COUNTEREXAMPLE反例与失效边界Counterexample and limits
VERIFICATION验证动作与证据Validation and evidence
OWNER责任人Accountable owner
VERSION版本与来源Version and source
EXPIRY有效期与复核日Expiry and review
WRITEDENIED UNTIL QA

冲突不是结案备注,而是规则更新的受控输入Conflict is a governed rule-update input—not a closing note

01 · RECORD不可变保存AI输出、工程师判断、证据差异Preserve AI, human and evidence delta
02 · REVIEW试点期每周复盘;重大安全分歧即时升级Weekly pilot review; immediate safety escalation
03 · PROPOSE形成评测集、规则或知识变更候选Propose eval, rule or knowledge change
04 · RELEASEQA批准、版本发布、回归监测;失败则回滚QA, versioned release, regression and rollback
闭环指标:冲突率、人工覆盖原因分布、同类冲突复发率、从冲突到裁决时长、规则更新后回归通过率。稳定期复盘频率由试点数据决定,不把“每周”固化成无依据制度。Track conflict rate, override reasons, recurrence, time to decision and post-change regression. Steady-state cadence is calibrated from pilot data.

记忆治理:复用次数不能替代版本、时效与适用性Memory governance: reuse count cannot replace version, age and applicability

weight = evidence_quality × exp(−ln2 × age_days / half_life) × applicability × version_match
CANDIDATEVERIFIEDACTIVESTALERETIRED
设备改造、配方/材料变更、传感器校准或SOP修订触发版本失配;version_match=0时立即排除,不因“被复用过”继续生效。30/60/180天半衰期仅为试点治理默认值,分别对应维修经验、配方/材料经验、正式SOP,必须由业务Owner基于复发周期校准。Retrofit, recipe/material change, calibration or SOP revision can set version_match=0 and immediately exclude a record regardless of reuse. The 30/60/180-day half-lives are pilot defaults for maintenance, recipe/material and approved SOP memory, pending owner calibration.

面对不能回答或不能执行的请求:拒绝后给出可继续的正确路径When a request cannot be answered or executed: refuse, then provide a safe path forward

09

思考作业三 · 交付REFLECTION 03 · DELIVERY

最大的机会是证据基础设施,最大的风险是规模化传播错误The largest opportunity is evidence infrastructure; the largest risk is scaled error

OPPORTUNITY

跨系统证据组织与经验复用Cross-system evidence and experience reuse

把MES、SCADA、设备日志、检测、SOP和沟通中的分散信息,对齐为带来源、时间、版本和责任人的证据包,让工程师把时间用于判断,而不是找资料。Align dispersed MES, SCADA, log, inspection, SOP and communication data into a traceable package so engineers spend time judging—not searching.

RISK

把概率推断包装成确定结论Probabilistic inference presented as certainty

多Agent、自动化与知识沉淀可能让同源偏差形成错误共识,并以更快速度、更高权威感扩散。控制手段是来源、反证、人工门槛、只读权限和审核后学习。Multi-agent workflows can turn correlated bias into false consensus and spread it faster. Controls are provenance, challenge evidence, human gates, read-only access and QA-approved learning.

多Agent增量ROI情景计算器Multi-agent incremental ROI scenario calculator

情景值 · 非实测SCENARIO · NOT MEASURED

月净增量价值Monthly incremental value

—

公式同时扣除Agent准备、人工复核、模型系统与误报处理成本。所有默认值仅用于展示测量结构。The formula deducts agent preparation, human review, system/model and false-alert costs. Defaults only demonstrate the measurement structure.

每一个ROI变量都必须能回到数据源与责任人Every ROI variable must trace to a source and owner

变量Variable操作定义Operational definition首选数据源Preferred source校准责任人Owner当前状态Status
N按统一口径进入调查的月案例数,去重后计数Deduplicated monthly investigation cases under one definitionMES / QMS / incident ticket业务OwnerBusiness ownerNOT MEASURED
T_before从责任人接单到首份可复核证据清单Assignment to first reviewable evidence list工单时间戳+两周时间研究Ticket timestamps + two-week time study产线工程师Line engineerPENDING BASELINE
T_agent / T_review系统执行时长与人工复核时长分开记录System run and human review measured separately不可变运行轨迹+复核事件Immutable traces + review events产品+质量Product + QualitySHADOW REQUIRED
C_hour含福利与间接成本的财务批准口径Finance-approved loaded labor rate财务成本口径Finance cost standard财务FinancePENDING FINANCE
C_model_infra模型、算力、存储、运维与审计留存的全成本Model, compute, storage, ops and audit retention账单+容量预算Billing + capacity planIT / FinOpsSCENARIO DEFAULT
F / T_false影子期误报数及其实际处理时长Shadow false alerts and actual handling time盲评标签+处理事件Blind labels + handling events质量OwnerQuality ownerSHADOW REQUIRED
停机/质量收益Downtime/quality benefit需有同期对照或可信反事实;当前公式明确排除Requires control/counterfactual; explicitly excluded nowMES / OEE / QMS生产+财务Production + FinanceEXCLUDED
测量纪律:至少2周或30例冻结Before;目标值不能冒充基线。只有在同口径Before/After、完整成本、误报负担和财务复核同时具备时才报告ROI。停机减少和质量提升没有反事实证据前均记为0。Freeze Before for at least two weeks or 30 cases. Targets are not baselines. Report ROI only with comparable Before/After, full costs, false-alert burden and finance review. Downtime and quality benefit remain zero without counterfactual evidence.

提交自检 · 40项通过Submission audit · 40 checks passed

下载中心Download center

课程来源|AI小项目实战-进阶(20—23集)Course · Advanced AI project practice, episodes 20–23CrewAI · FastAPI · RAG · Multi-Agent codingUCI SECOM Dataset1567 samples · 590 features · missing values · CC BY 4.0
企业事实状态Enterprise fact statusTCL岗位、平台、基线、接口、盲测与ROI均待授权核验TCL roles, stack, baselines, interfaces, blind tests and ROI remain pending authorization