DAY 14 · REAL CASE · VERIFIED SOURCES

从看见到创作FROM SEEING多模态视觉生产链TO CREATING

以武汉大学老图书馆真实照片为唯一案例,完成AI识图、事实核验、宣传文案、校园海报与24秒创意短片。不是展示概念,而是展示一条可以检查、可以追溯、可以交付的完整证据链。

One licensed photograph of Wuhan University's Old Library drives the complete chain: visual interpretation, evidence auditing, promotional copy, poster design and a 24-second creative film.

武汉大学老图书馆真实照片
建筑主体 · VISIBLE
成熟树木 · VISIBLE
具体年代 · VERIFY
00 / 提交核验00 / COMPLIANCE

四项要求,不靠口头声明

FOUR REQUIREMENTS, EACH WITH VISIBLE EVIDENCE

每张卡片严格对应老师原始提交要求。点击即可跳到原图、问答、三栏分析、海报、视频、提示词和规定字数说明所在位置。

Each card maps the original brief to visible evidence. Open any card to reach the source image, Q&A, three-column audit, poster, film, prompts and required-length reflections.

作业 01ASSIGNMENT 01
证据完整COMPLETE

图文转换:识图、宣传稿与错误订正

IMAGE-TO-COPY: VISION, CAMPAIGN TEXT & CORRECTION

80—150字文案100字以上评价
原图与AI结果记录可见
Original image and AI record visible
标题、口号与最终宣传稿完整
Title, slogan and final copy included
三处风险逐项订正
Three risks corrected item by item
作业 02ASSIGNMENT 02
证据完整COMPLETE

能力边界:AI会不会“看错”

CAPABILITY BOUNDARY: WHEN VISION OVERREACHES

连续3问事实/推断/猜测100字以上结论
三轮问题与完整回答
Three prompts and complete answers
三栏证据分类
Three-column evidence classification
结论明确标注字数达标
Conclusion clearly meets length rule
作业 03ASSIGNMENT 03
证据完整COMPLETE

综合创作:AI校园主题海报

COMPOSITE CREATION: CAMPUS POSTER

2种AI工具关键提示词150字以上说明
海报成品可查看与下载
Poster can be viewed and downloaded
两种工具分工与提示词
Two-tool roles and prompts
创作说明独立成段
Standalone creative rationale
作业 04ASSIGNMENT 04
证据完整COMPLETE

综合创作:24秒AI创意短片

COMPOSITE CREATION: 24-SECOND AI FILM

15—30秒分镜/脚本150字以上复盘
六镜头成片直接播放
Six-scene film plays in the page
工具、主提示词与运动约束
Tools, master prompt and motion constraints
视频及分镜原图可下载
Film and full storyboard frames downloadable
01 / SOURCE

真实素材,先于AI结论

SOURCE BEFORE CONCLUSION

原图、作者、日期、许可、文件校验值和官方事实被独立保存,后续任何AI输出都不能覆盖这份原始证据。

The original image, author, date, license, checksum and official facts are fixed before any AI interpretation.

武汉大学老图书馆原始照片
Whu old library 1.JPG · 3872×2592 · Howchou · 2012-01-04
SOURCE TRACE

素材与核验档案

来源Wikimedia Commons
许可CC BY-SA 3.0
原文件2.32 MB · 直接摄影图像
校验值c4baee90a18a48c24336522a5b696aa3a805bdf99181b0109c360305412e768b
官方事实武汉大学图书馆官方资料确认:老图书馆大楼于1935年启用。
边界具体名称、年份和文保身份不能仅从像素中直接读出。

图片改编成品继续采用 CC BY-SA 3.0,并保留摄影作者和来源。

02 / VISION

AI看见了什么,又猜了什么?

WHAT DID AI SEE — AND ASSUME?

三次连续提问逐步收紧约束,观察模型是否把视觉事实、合理推断和既有知识混为一谈。

Three consecutive prompts progressively separate visible evidence, inference and unsupported certainty.

PROMPT 01只列出图片中能够直接看见的物体和视觉特征,不判断具体地点,不使用图片以外的知识。
AI RESPONSE画面中可见一座多层建筑,具有灰绿色瓦屋顶、起翘屋檐、中央高起部分、窗户、立柱和装饰屋脊。建筑周围有成熟树木,前景有道路、台阶、矮墙、栏杆、提示牌、红色覆盖物和一名距离较远的人。
PROMPT 02根据画面判断,这可能是什么类型的场景?请给出依据并标明置信度。
AI RESPONSE · 75%可能是校园、文化机构或纪念性公共建筑。依据是建筑体量、公共入口、成熟绿化和较安静的环境;画面中没有清晰可读的校名或馆名,因此不能排除其他公共建筑。
PROMPT 03能否仅根据图片确定建筑名称、地点、建成年份和用途?哪些能确定,哪些不能确定?
AI RESPONSE仅根据图片不能可靠确定具体名称、城市、建成年份和当前用途。可以确定其视觉特征,可以推断它可能位于校园或文化场所;具体身份必须查看图片出处或权威介绍。
01

事实 FACT

  • 大型多层建筑
  • 灰绿色瓦屋顶
  • 中央塔楼与装饰屋脊
  • 成熟树木、道路、台阶
  • 晴朗暖光环境
02

推断 INFERENCE

  • 可能是校园或文化机构
  • 可能承担公共功能
  • 具有传统中式视觉表达
  • 可能位于较高地势
  • 置信度必须保留
03

猜测 GUESS

  • 一定是武汉大学老图书馆
  • 1935年启用
  • 具体设计者
  • 文物保护等级
  • 当前内部用途
100字以上结论100+ CHARACTER CONCLUSION

AI会“看错”吗?

CAN AI VISION BE WRONG?

计算中

连续三问显示,AI最可靠的能力是描述画面中直接可见的颜色、形状、物体与空间关系;当问题转向场景类型时,回答已经进入概率推断;当问题要求具体建筑名称、地点、年份和用途时,如果没有来源支撑,就可能调用训练记忆并形成没有依据的确定语气。因此,AI的“看见”并不等同于摄像机记录,而是像素观察、模式匹配、常识推断和既有知识的混合结果。正确做法不是禁止AI推断,而是要求它标明依据、置信度和不可确定项,再由人工使用图片来源页与权威资料核验。流畅表达不是证据,高置信度也不等于事实,最终发布责任必须由人承担。

Across three questions, AI is most reliable when describing visible colour, shape, objects and spatial relations. Scene type already involves probabilistic inference; exact identity, location, date and use can trigger prior knowledge without sufficient evidence. AI “seeing” is therefore a mixture of pixel observation, pattern matching, common-sense inference and learned memory. The correct workflow is to require evidence, confidence and explicit unknowns, then verify claims through source pages and authoritative records. Fluency is not evidence, confidence is not fact, and publication responsibility remains human.

查看可独立提交的AI问答结果截图AI识图与三轮问答结果记录
03 / COPY

从识图结果到宣传语言

FROM OBSERVATION TO PUBLIC COPY

文案严格控制在80—150字,将可核验事实与画面感结合,不用虚构故事制造“历史感”。

The final copy combines verified facts with visual description, without inventing a historical narrative.

COPY SYSTEM

标题

山水有书声|武汉大学老图书馆

口号

登珞珈之巅,读百年书声。

核验事实

老图书馆大楼于1935年启用;老馆及周围建筑群被列入第五批全国重点文物保护单位。

宣传文案 · 108字 · 达标
沿珞珈山拾级而上,武汉大学老图书馆静立于狮子山顶。1935年启用的建筑,将知识殿堂与山水校园连接在一起。飞檐、屋脊与层层台阶记录着一代代求学者的足迹。走近它,不只是观看一座建筑,更是在阅读一所大学跨越岁月的文化记忆。
登珞珈之巅,读百年书声。
错误订正记录CORRECTION RECORD

AI没有“硬认错”,但三处表达必须降级

NO HARD MISIDENTIFICATION, BUT THREE CLAIMS REQUIRE DOWNGRADING

作业一必交证据REQUIRED EVIDENCE
AI表达AI STATEMENT问题RISK人工订正HUMAN CORRECTION
“可能是校园、文化机构或纪念性公共建筑”“Possibly a campus, cultural institution or commemorative public building.”这是基于体量和环境的场景推断,不是画面事实。This is a scene inference based on scale and context, not a visible fact.保留“可能”与75%置信度,不能直接写进最终宣传稿。Keep “possibly” and 75% confidence; do not publish it as fact.
“具有传统中式视觉表达”“Shows traditional Chinese visual language.”只能描述屋顶、飞檐等视觉特征,不能代替完整建筑风格鉴定。Visible roof and eave features do not constitute a full architectural classification.改为“画面可见灰绿色瓦顶、起翘屋檐与装饰屋脊”。Revise to the directly observable tiled roof, upturned eaves and decorative ridges.
“可能位于较高地势”“Possibly located on elevated ground.”单一机位和透视关系不足以确认海拔或地势。One camera angle cannot establish elevation or terrain.降级为猜测;地点、1935年与文保身份均改由来源页和官方资料核验。Downgrade to guesswork; verify location, 1935 and heritage status through source and official records.
100字以上评价100+ CHARACTER EVALUATION

识图结果评价

EVALUATION OF THE VISION RESULT

计算中

这次识图最有价值的地方,不是AI说出了建筑名称,而是它在限制条件下把可见物体、场景推断和无法确认的信息分开。第一轮对屋顶、飞檐、树木、台阶和道路的描述基本准确,可以作为后续文案的视觉素材;第二轮提出校园或文化机构的可能性,但这只能作为带置信度的推断;第三轮明确承认名称、地点、年份和用途不能仅凭图片确定,避免了把训练记忆伪装成视觉证据。人工订正进一步将“传统中式表达”和“较高地势”降级处理,并用图片来源与武汉大学官方资料补足身份和1935年启用信息。最终宣传稿因此既保留画面感,也没有把猜测写成事实。

The strongest part of the experiment is not naming the building, but separating visible objects, scene inference and unverifiable information. The first response accurately supplies visual material; the second remains a confidence-qualified inference; the third explicitly acknowledges that identity, place, date and use require external evidence. Human correction further downgrades architectural-style and terrain claims, while the source page and Wuhan University records verify identity and the 1935 opening date. The final copy therefore remains evocative without turning guesses into facts.

04 / POSTER

真实照片承担事实,AI视觉承担氛围

REAL PHOTO, AI-ASSISTED ATMOSPHERE

海报没有让AI重建一座“看起来更壮观”的假建筑,而是保留真实主体,用AI完成色彩、蓝图线和文化纹理探索,最后以精确排版控制全部文字。

The real building remains the factual anchor. AI contributes restrained atmosphere and texture; exact typography remains deterministic.

山水有书声校园文化宣传海报
原始照片AI辅助视觉背景原始照片AI辅助背景
AI TOOL 01ChatGPT 多模态理解

识图、事实边界、文案、信息层级与提示词。

AI TOOL 02OpenAI 图像生成

建筑照片视觉延展、蓝图线、纸张纹理与光线。

工具一|查看信息规划提示词TOOL 01 | VIEW INFORMATION-PLANNING PROMPT
你是一名校园文化海报策划与事实编辑。基于已核验的武汉大学老图书馆原图、来源信息和宣传文案,输出海报的信息层级:主标题、口号、核心事实、视觉焦点、署名与许可;说明哪些内容来自图片,哪些来自官方资料。不得虚构人物故事、设计者信息或建筑用途。输出为“信息层级 / 视觉位置 / 事实依据 / 风险检查”四列表。
Act as a campus-culture poster strategist and fact editor. Using the verified source image, provenance and campaign copy, define the hierarchy for title, slogan, verified fact, visual focus, credit and license. Separate image evidence from official-source facts. Do not invent personal stories, architects or building uses. Output four columns: hierarchy, placement, evidence and risk check.
工具二|查看图像生成提示词TOOL 02 | VIEW IMAGE-GENERATION PROMPT
以输入的武汉大学老图书馆真实照片为唯一建筑主体,保持屋顶、塔楼、檐口和整体比例可辨识;制作作品集级竖版校园文化海报背景。使用深墨蓝、灰青、纸张米白和克制暖金色,加入轻量建筑蓝图线、山形等高线与书页纹理。顶部保留标题安全区,底部保留来源说明区。禁止文字、校徽、水印、人物、虚构建筑细节和廉价古风特效。
Use the supplied real photograph of Wuhan University Old Library as the sole architectural subject. Preserve the identifiable roof, tower, eaves and proportions. Create a portfolio-grade vertical campus-culture poster background in deep indigo, grey-cyan, paper white and restrained warm gold, with subtle blueprint lines, topographic contours and page texture. Keep title and source safe areas. No text, emblem, watermark, people, invented architecture or cheap pseudo-classical effects.
150字以上创作说明150+ CHARACTER CREATIVE RATIONALE

两种AI能力如何形成一张可交付海报

HOW TWO AI CAPABILITIES PRODUCED A DELIVERABLE POSTER

计算中

海报坚持“真实照片承担事实,AI视觉承担氛围,人工排版承担信息准确”的分工。第一种工具ChatGPT多模态理解先整理原图中可见的建筑、树木和台阶,区分图片事实与外部核验信息,再规划主标题、口号、1935年核心事实、图片署名和许可的阅读层级。第二种工具OpenAI图像生成以原图为建筑身份参考,只扩展深墨蓝、暖金、蓝图线、山形等高线和纸张纹理,不负责生成任何文字。最终版由可控排版完成中文标题、年份、来源和CC BY-SA 3.0许可,避免AI文字乱码。构图从山脚、建筑到天空建立向上的阅读方向,对应“拾级而上”的叙事;大面积留白保证标题清晰,暖金光线集中视线,底部完整保留摄影作者和来源。这样既利用了AI的理解与视觉生成能力,又没有把生成效果误当成历史事实。

The poster follows a strict division of labour: the real photograph carries fact, AI carries atmosphere, and deterministic typesetting carries information accuracy. ChatGPT multimodal analysis first separates visible architecture, trees and steps from externally verified facts, then plans the title, slogan, 1935 fact, credit and license hierarchy. OpenAI image generation uses the photograph as an identity reference and extends only colour, blueprint lines, topographic contours and paper texture; it generates no final text. Controlled layout then places the Chinese title, date, source and CC BY-SA 3.0 notice without AI text errors. The upward composition supports the ascent narrative while whitespace protects readability and authorship remains explicit.

05 / FILM

24秒,六个镜头完成一次记忆抵达

SIX SHOTS. ONE JOURNEY THROUGH MEMORY.

重制短片采用六个独立画面、五种转场与原创声音设计,从建筑蓝图、黎明山路、学生登阶、屋檐细节、书页隐喻推进到蓝调时刻收束,形成完整的视觉叙事。

The remastered film uses six distinct scenes, five transitions and original sound design—from archival blueprint and dawn ascent to architectural detail, the book metaphor and a blue-hour finale.

六个镜头具有独立构图和叙事功能,并分别使用推近、拉远、向上跟随、横向细节扫描和光感叠化。音轨由低频氛围、风声层与五个转场提示音原创合成;成片为24.000秒、1080×1920、30fps。

Six scenes have independent compositions and narrative roles, using push-in, pull-back, upward tracking, lateral detail scan and light-based dissolves. The original track combines low-frequency ambience, wind texture and five transition cues. Final output: 24.000 seconds, 1080×1920, 30 fps.

使用工具说明TOOLS & RESPONSIBILITIES

从策划到成片的实际分工

ACTUAL WORKFLOW FROM STORY TO FINAL FILM

ChatGPT

拆解24秒叙事弧、核验信息边界、规划六个镜头功能、撰写字幕和负面约束。

Built the 24-second arc, audited factual boundaries, assigned six scene roles, and drafted captions and negative constraints.

OpenAI Image Generation

以真实建筑照片为身份参考,分别生成蓝图、黎明、学生、屋檐、书页和蓝调结尾六张独立画面。

Used the real building as an identity reference to generate six distinct blueprint, dawn, student, eaves, page and blue-hour scenes.

FFmpeg

完成推拉运镜、五种转场、字幕烧录、原创声音合成、H.264/AAC编码与24.000秒精确裁切。

Created camera movement, five transitions, subtitle burn-in, original sound design, H.264/AAC encoding and exact 24.000-second timing.

视频提示词说明VIDEO PROMPT DOCUMENTATION

主提示词与运动约束

MASTER PROMPT & MOTION CONSTRAINTS

六镜头主提示词SIX-SCENE MASTER PROMPT
以武汉大学老图书馆真实照片为建筑身份与结构参考,制作六个具有独立叙事作用的9:16电影级镜头:①档案蓝图与珞珈山等高线显影;②黎明山路与石阶引向建筑;③一名学生携书从背后登阶;④灰绿色瓦片、飞檐与中央塔楼特写;⑤翻动书页化为建筑轮廓与知识光路;⑥蓝调时刻的建筑、暖窗和石阶完成收束。保持中央塔楼、层叠屋顶、飞檐和整体比例可辨识。深墨蓝、灰青与克制暖金,保留字幕安全区。禁止文字、水印、额外塔楼、建筑变形、多人、廉价赛博朋克HUD、玄幻爆炸和过度粒子。
Use the real photograph of Wuhan University Old Library as the identity and structure reference. Create six independent 9:16 cinematic scenes: archival blueprint reveal; dawn path and steps; one student ascending with books; roof, eaves and tower detail; turning pages transforming into the building; and a blue-hour finale with warm windows and lit steps. Preserve the central tower, layered roofs, upturned eaves and proportions. Use indigo, grey-cyan and restrained warm gold with subtitle-safe space. No text, watermark, extra towers, warped architecture, crowds, cheap cyberpunk HUD, fantasy explosions or excessive particles.
剪辑与运动约束EDITING & MOTION CONSTRAINTS
每镜头4.5秒,0.6秒转场重叠;依次使用缓慢推近、拉远显景、向上跟随、横向细节扫描、缓慢拉远和结尾推近。镜头不得抖动,不允许主体形变或闪烁。五个转场点使用溶解、向上平滑、淡化、径向和淡黑;字幕位于底部安全区,避开人物与建筑主体。
Each scene runs 4.5 seconds with 0.6-second overlap. Use slow push-in, pull-back reveal, upward tracking, lateral detail scan, slow pull-back and final push-in. No shake, subject deformation or flicker. Use dissolve, smooth-up, fade, radial and fade-to-black transitions. Keep subtitles in the lower safe area away from people and architecture.
150字以上复盘150+ CHARACTER REVIEW

为什么这次不是“单图平移”

WHY THIS IS NO LONGER A SINGLE-IMAGE PAN

计算中

重制短片不再把一张照片反复放大和移动,而是围绕“显影—抵达—观看—理解—延续—收束”设计六个独立镜头。蓝图镜头负责提出记忆主题,黎明远景建立珞珈山空间,学生登阶把历史建筑与当代学习者连接起来,屋檐特写承载1935年启用这一核验事实,书页化形表达知识传承,蓝调结尾用暖窗和石阶光路收束口号。剪辑分别使用推近、拉远、向上跟随、横向扫描与光感叠化,避免六张静帧变成机械幻灯片。声音由低频氛围、风声层和五个转场提示音原创合成,没有使用未授权音乐或克隆声音。最终成片严格控制为24.000秒、1080×1920、30fps,并对六个时间点逐帧检查人物、建筑、字幕和转场,确保信息准确、画面稳定且符合15—30秒要求。

The remastered film replaces repeated zooming on one photograph with six independently designed scenes following a reveal-arrival-observation-understanding-continuation-resolution arc. Blueprint introduces memory; dawn establishes place; the student connects history with the present; the eaves carry the verified 1935 fact; pages express knowledge transmission; and blue hour resolves the slogan. Camera movement and five transitions prevent the sequence from becoming a static slideshow. The original soundtrack uses low-frequency ambience, wind texture and transition cues without unauthorised music or voice cloning. The final 24.000-second, 1080×1920, 30 fps film was checked at six timestamps for people, architecture, captions and transitions.

06 / DELIVERY

全部素材,页面内可预览也可下载

PREVIEW AND DOWNLOAD EVERY ASSET

原图、AI结果、海报、视频封面与六张完整分镜均提供独立下载;视频、PDF和文字证据提供明确文件名,单文件HTML内无需依赖外部文件夹。

The source image, AI record, poster, film cover and all six full storyboard frames are independently downloadable. The film, PDF and evidence text retain explicit filenames even inside the self-contained HTML.

图片与分镜下载

IMAGE & STORYBOARD DOWNLOADS

成品文件下载FINAL FILE DOWNLOADS

视频、提交汇总 PDF 与完整文字证据

FILM, FINAL SUBMISSION PDF & COMPLETE TEXT EVIDENCE

作业一ASSIGNMENT 01

原图、结果记录、108字文案、订正、100字以上评价。

Source, AI record, 108-character copy, corrections and 100+ evaluation.

作业二ASSIGNMENT 02

连续三问、完整回答、三栏整理、100字以上结论。

Three questions, complete answers, three columns and 100+ conclusion.

作业三ASSIGNMENT 03

海报、两种工具、两组提示词、150字以上说明。

Poster, two tools, two prompt sets and 150+ rationale.

作业四ASSIGNMENT 04

24秒成片、六镜头分镜、工具与提示词、150字以上复盘。

24-second film, six-scene storyboard, tools, prompts and 150+ review.