Transcript API Guide财报电话会议文字稿 API 接入指南

Build with a Free Earnings Transcript API如何接入免费财报电话会议文字稿 API

Compare transcript coverage, licensing, structure, and limits before turning quarterly calls into searchable research data.

先比较文字稿覆盖范围、使用许可、内容结构与请求限制,
再将季度财报电话会议内容转化为可检索的研究数据。

Hand-drawn workflow for finding, parsing, storing, and searching earnings call transcripts

Core summary核心摘要

A useful earnings transcript API returns more than a block of text. It preserves the event identity, company, fiscal period, call time, publication and revision times, prepared remarks, question-and-answer sequence, speaker names and roles, and a source record that supports quotation. For research and AI applications, every extracted claim should remain traceable to a specific speaker turn in a specific transcript version.

可靠的财报电话会议文字稿 API 不应只返回一整段文本。它需要保留活动身份、公司、财务报告期、会议时间、发布时间与修订时间,并区分管理层陈述和问答环节,记录发言人姓名、职务及原始来源。用于研究或 AI 应用时,每一项提取结论都应能够追溯到某个文字稿版本中的具体发言片段。

What the API returns

Look for prepared remarks, Q&A, speaker names, company, fiscal quarter, and publication metadata.

What “free” means

Free may mean a small quota, delayed access, limited history, or non-commercial terms.

Production rule

Store raw text, normalized segments, source URLs, and retrieval times so quotations remain traceable.

接口应返回哪些内容

应包含管理层陈述、问答环节和发言人姓名,以及公司、财季和发布时间等元数据。

“免费”的实际含义

免费服务可能仅提供少量配额、延迟数据或有限的历史数据,也可能只允许非商业用途。

生产环境的数据留存原则

保存原始文本、标准化内容片段、来源网址和获取时间,确保每处引文均可追溯。

Understand transcript coverage before choosing an API选择接口前先明确文字稿覆盖范围

Coverage differs by exchange, company size, language, event type, and age. Test whether the dataset includes special calls, investor days, guidance updates, merger calls, and corrected versions—not only scheduled quarterly results. Also distinguish an official transcript, a vendor-edited transcript, automated speech recognition, and an audio-only record; accuracy, publication speed, and quotation rights differ.

不同接口在交易所、公司规模、语言、活动类型和历史数据跨度方面差异明显。除定期季度业绩会外,还要测试数据集是否涵盖特别电话会议、投资者日、业绩指引更新、并购说明会和修订版文字稿。同时应区分公司官方文字稿、供应商编辑稿、自动语音识别稿和仅提供音频的记录,因为它们在准确性、发布时间和引用权限方面并不相同。

Source type来源类型What it is useful for适合用途What must remain visible必须保留的信息Do not assume不能默认
Company-published transcript公司发布的文字稿Durable quotation and review when the company identifies it as the official record.公司明确其为正式记录时,适合长期引用与复核。Publisher, event, publication time, revision notice, source URL, and any omitted portions.发布方、活动、发布时间、修订说明、来源网址及是否存在删减。That every company publishes one, or that the text is verbatim rather than edited.不能默认所有公司都会发布,也不能默认文本完全逐字且未经编辑。
Vendor-edited transcript供应商编辑稿Structured speaker turns, normalized names, punctuation, and broad historical retrieval.适合获取结构化发言轮次、规范姓名、标点和较广的历史覆盖。Vendor, edition status, edit time, speaker-resolution method, license, and correction history.供应商、版本状态、编辑时间、发言人识别方法、许可与更正历史。That edited wording exactly matches the live audio or can be redistributed freely.不能默认编辑后的措辞与现场音频完全一致,也不能默认可以自由再分发。
Automated speech-recognition draft自动语音识别初稿Fast monitoring, event detection, and locating passages for later verification.适合快速监控、识别事件,以及定位需要后续复核的段落。Preliminary status, model or provider, confidence where available, audio offsets, and replacement version.初稿状态、模型或供应商、可用的置信度、音频位置与后续替代版本。That plausible names, numbers, negations, and speaker labels are correct.不能因为文字读起来通顺,就默认姓名、数字、否定词和发言人标签正确。
Audio or webcast only仅音频或网络直播Primary evidence when access and retention rights permit independent transcription.在访问和留存许可允许时,可作为独立转写的一手证据。Event time, language, audio duration, access terms, retrieval time, and segment offsets.活动时间、语言、音频时长、访问条款、获取时间和片段位置。That streaming access permits downloading, storage, transcription, or public quotation.不能默认可以下载、保存、转写或公开引用。
Company, event, and fiscal identity

Test large, small, delisted, renamed, and dual-listed companies. Verify stable company and event IDs, event type, fiscal year and quarter, call start time, timezone, and whether the event actually corresponds to the reported period.

Transcript structure

Require prepared remarks and Q&A boundaries, speaker turns in order, executive and analyst roles, operator instructions, and paragraph or sentence offsets. Speaker labels should remain separate from the spoken text.

Publication speed and status

Measure call-to-first-text and call-to-final-transcript delay. A preliminary automated transcript may support same-day monitoring, while an edited or official version may arrive later and replace uncertain wording.

Language and transcription quality

Check the original language, translations, punctuation, numbers, company and product names, acronyms, and speaker attribution. Test low-quality audio, overlapping speakers, and calls containing several languages.

Revision and source handling

Check whether corrections replace records or create versions. Preserve event time, publication time, update time, transcript status, source URL, and a checksum so citations do not silently point to changed text.

公司、活动与财务期间身份

分别用大型、小型、已退市、已更名和双重上市公司的数据进行测试。核对稳定的公司与活动 ID、活动类型、财年和财季、会议开始时间、时区,并确认会议确实对应所标注的财务报告期。

文字稿结构

应明确划分管理层陈述和问答环节,按顺序记录每次发言,区分公司高管、分析师和主持人角色,并保留段落或句子位置。发言人标签应与其发言正文分开存储。

发布速度与版本状态

分别测量从会议开始到首次文本、以及到最终文字稿可用的时间。自动生成的初稿可以支持当日监控,经过编辑或由公司发布的正式版本通常更晚,但会修正不确定措辞。

语言与转写质量

检查原始语言、翻译、标点、数字、公司和产品名称、缩写及发言人归属,并测试音质较差、多人重叠发言和一场会议使用多种语言的情况。

修订版本与来源处理

确认更正内容会覆盖原记录还是生成新版本,并保留会议时间、发布时间、更新时间、文字稿状态、来源网址和校验和,避免引文悄然指向已经变化的文本。

Evaluate a free earnings transcript API with real calls使用真实业绩会数据评估免费财报电话会议文字稿 API

Test测试项What to inspect检查内容Failure signal风险信号Engineering response工程应对
Structure内容结构Prepared remarks, Q&A, speaker role, sequence, and timestamps.管理层陈述、问答环节、发言人角色、发言顺序和时间戳。One unlabelled text block.仅返回一整段不含章节或发言人标记的文本。Create segments but preserve the raw response.将文本拆分为结构化片段,同时保留原始响应。
Identity标识体系Stable company, event, quarter, and speaker identifiers.稳定的公司、业绩会、财季和发言人标识符。Ticker-only matching.仅凭股票代码匹配记录。Maintain an entity mapping table.维护实体映射表。
Text quality文本质量Numbers, names, acronyms, punctuation, omitted audio, and confidence markers.数字、名称、缩写、标点、未识别音频和置信度标记。Plausible but incorrect words or figures.文字看似通顺,但关键名称或数字转写错误。Compare difficult samples with audio or an official source.选择高难度样本,与音频或公司官方来源逐项核对。
Versioning版本管理Preliminary, edited, official, corrected, and translated versions.初稿、编辑稿、官方稿、更正版与翻译版本。Text changes under the same record ID.同一记录 ID 下的文字发生变化,却没有版本记录。Store checksums and retain prior versions.保存校验和,并保留此前版本。
Limits请求限制Daily quota, burst rules, pagination, and retry headers.每日请求配额、突发请求上限、分页机制和重试响应头。Undocumented 429 responses.未说明何时会返回 429 状态码。Queue, cache, back off, and monitor quota.使用请求队列和缓存,执行退避重试,并监控配额用量。
Rights使用许可Storage, analysis, quotation, redistribution, and attribution terms.有关存储、分析、引用、再分发和署名的许可条款。A free key with unclear license.API 密钥虽免费,但许可条款不明确。Block launch until rights are documented.使用权条款明确并留档前,不得上线。

Build a reliable transcript ingestion workflow构建可靠的业绩会文字稿采集流程

Discover the event before requesting text

Resolve the company and earnings event, then query by stable event ID where possible. Confirm fiscal period, call time, timezone, event type, transcript language, and expected publication status.

Save an immutable source snapshot

Store the raw payload, request, source URL, retrieval time, status, version, and checksum before transformation. If audio is available and licensed, retain its reference for transcription review.

Normalize with evidence

Record sections and speaker turns in sequence, keeping names, roles, offsets, and source text. Treat speaker resolution as a separate field so a later identity correction does not rewrite the quotation.

Index for quotation-aware research

Chunk by speaker turn or coherent topic rather than arbitrary character count. Attach company, event, quarter, date, section, speaker, role, transcript version, and source offsets to every chunk.

Monitor corrections and generated outputs

Detect revisions with checksums, retain prior versions, rebuild affected indexes, and flag summaries whose cited source changed. Generated claims should quote sparingly and link back to the underlying turn.

Run a quotation and extraction acceptance test

Choose calls with prepared remarks, dense Q&A, unclear speaker handoffs, corrected numbers, product names, and forward-looking guidance. For each test, verify the quoted words, speaker, section, fiscal period, numeric unit, and source offsets. Score retrieval and extraction separately: finding the right passage does not prove that the generated claim preserved its meaning.

获取文本前先确认活动身份

先解决公司与财报活动映射,条件允许时按稳定的活动 ID 查询。核对财务报告期、会议时间、时区、活动类型、文字稿语言和预期发布状态。

保存不可变的来源快照

转换前保存原始响应、请求参数、来源网址、获取时间、状态、版本和校验和。如果许可允许且接口提供音频,还应保留音频引用,以便复核转写质量。

标准化并保留证据

按顺序记录章节和发言轮次,同时保存姓名、职务、文本位置和来源正文。发言人身份应作为独立字段处理,后续纠正身份时不应改写原始引文。

建立支持引文的研究索引

按完整发言轮次或语义连贯的主题分段,而不是机械按字符数切分。每个片段应附上公司、活动、财季、日期、章节、发言人、角色、文字稿版本和原文位置。

监控修订和已生成内容

使用校验和检测修订,保留先前版本,重建受影响的索引,并标记来源已经变化的摘要。生成结论应克制引用,并链接回对应发言片段。

执行引文与信息抽取验收

测试集应覆盖管理层陈述、密集问答、发言人切换不清、数字更正、产品名称和前瞻指引。逐条核对引文原文、发言人、所在章节、财务期间、数值单位和原文位置,并把检索准确性与抽取准确性分开评分:找到正确段落,并不代表生成结论保留了原意。

Consider a hypothetical call in which the CEO says during prepared remarks, “We expect next-quarter revenue growth of 10% to 12%.” In Q&A, an analyst asks whether that range includes a recently acquired business, and the CFO answers that the outlook includes two months of acquired revenue but excludes a possible currency headwind. A summary that stores only “guidance: 10%–12%” loses the scope conditions that make the number meaningful.

假设某场业绩会上,CEO 在管理层陈述中表示:“预计下一季度收入增长 10% 至 12%。”到了问答环节,分析师追问该区间是否包含最近收购的业务,CFO 回答称,指引已计入两个月的收购业务收入,但没有计入潜在汇率逆风。如果摘要只保存“收入指引为 10% 至 12%”,就会丢失决定该数字含义的范围条件。

Evidence unit证据单元Store separately需要分别保存Why separation matters为什么不能合并
Prepared-remarks statement管理层陈述Speaker, role, section, exact range, metric, target period, source offsets, and transcript version.发言人、职务、章节、准确区间、指标、目标期间、原文位置和文字稿版本。This is the initial guidance claim, not the complete set of assumptions.这是最初的指引结论,并不包含全部假设。
Analyst question分析师问题Analyst identity, affiliation, question text, referenced claim, and sequence.分析师身份、所属机构、问题原文、所指结论和发言顺序。A question is context, not a fact asserted by management.问题提供上下文,但不能当作管理层确认的事实。
Management clarification管理层补充说明Responding speaker, acquisition treatment, currency assumption, confidence or uncertainty, and link to the question.回答者、收购业务处理、汇率假设、确定或不确定程度,以及与问题的关联。These qualifiers change how the numerical range should be interpreted and compared.这些限定条件会改变数字区间的解释方式和可比性。
Derived research claim衍生研究结论Normalized metric, value range, period, scope flags, source turn IDs, extraction method, and review status.标准化指标、数值区间、期间、范围标记、来源发言 ID、抽取方法和复核状态。Readers can verify both the headline number and the conditions attached to it.读者可以同时核对核心数字及其附带条件。

Revision rule: when a final transcript corrects “10% to 12%” to “10% to 11%,” create a new transcript version, link the changed turn, and mark dependent summaries for review. Do not rewrite the old record in place. Historical monitoring may need to preserve what was available at each publication time, while current research should normally point to the latest verified version.

修订规则:如果最终版文字稿把“10% 至 12%”更正为“10% 至 11%”,应创建新的文字稿版本,关联发生变化的发言片段,并把依赖该引文的摘要标记为待复核,不能直接覆盖旧记录。历史监控可能需要保留每个发布时间点当时可见的内容;当前研究则通常应指向最新且已经核验的版本。

Use QVeris to discover transcript capabilities使用 QVeris 查找文字稿类能力

QVeris helps developers and agents find transcript, event-text, and corporate-call capabilities, inspect inputs and outputs, and make provider assumptions explicit before production use. It can simplify discovery and calling; the application still owns event identity, transcript versioning, speaker resolution, quotation rights, and the evidence attached to generated analysis.

QVeris 帮助开发者和智能体发现文字稿、活动文本和公司电话会议能力,核对输入与输出,并在用于生产环境前明确记录对供应商能力与行为的假设。它可以简化能力发现和调用,但活动身份、文字稿版本、发言人映射、引用权限以及生成分析所附带的证据,仍由应用方负责。

  • Search by capability, such as earnings transcripts, event text, speaker-labelled calls, or historical corporate events.
  • Inspect authentication, identifiers, required parameters, response fields, pagination, and request frequency limits before integration.
  • Keep a provider adapter and provenance trail even when discovery and calling are streamlined.
  • 按能力搜索,例如业绩会文字稿、活动实录、带发言人标签的电话会议记录或公司历史活动记录。
  • 接入前检查鉴权方式、各类标识符、必填参数、响应字段、分页机制和请求频率限制。
  • 即使查找和调用流程已经简化,仍应保留数据提供商适配层和完整的数据溯源链路。

FAQ常见问题

Is there a truly free API?

Yes, usually with quotas, delay, limited history, or restricted usage rights.

Can transcripts power AI summaries?

Yes when licensing permits. Preserve citations and label generated summaries.

What should I cache?

Cache raw responses and normalized segments; refresh metadata when corrections appear.

How should speakers be identified?

Keep the source label, normalized person ID, company or analyst affiliation, role, and confidence separately. Do not infer an identity from a surname alone when the quotation will be published.

Can preliminary transcripts be cited?

Only if the license permits and the preliminary status is visible. Prefer the final or official version for durable research, and recheck citations when a corrected transcript arrives.

What is the best way to chunk a transcript?

Start with speaker turns and section boundaries, then split only unusually long turns by topic while retaining source offsets. Arbitrary fixed-size chunks can separate a question from its answer.

是否有真正免费的 API?

有,但通常存在请求配额、数据延迟、历史数据覆盖有限或使用许可受限等条件。

财报电话会议文字稿可以用来生成 AI 摘要吗?

可以,但须获得相应许可。应保留引文的来源信息,并明确标注摘要由 AI 生成。

应缓存哪些数据?

应缓存原始响应和标准化后的内容片段;文字稿出现更正时,还应更新相关元数据。

应该如何识别发言人?

应分别保存来源中的发言人标签、标准化人员 ID、所属公司或分析机构、职务和识别置信度。引文需要公开时,不能仅凭姓氏推断具体身份。

初版文字稿可以引用吗?

只有在许可允许且明确标注“初稿”状态时才可引用。长期研究应优先采用最终版或官方版本;更正版发布后,还要重新核对已有引文。

文字稿应该如何切分?

优先按发言轮次和章节边界切分;只有单次发言过长时,才按主题继续拆分,并保留原文位置。机械使用固定长度切分,可能把问题与回答割裂。

External references外部参考链接