バグ/機能要求を報告

エクスポート形式

必要な形式で文字起こしをダウンロード。STT.aiは6つのエクスポート形式をサポートし、それぞれ異なるワークフローに最適化されています。

公開されているオーディオとビデオで動作します。DRM 保護されたコンテンツはサポートされていません。

アップグレード

Private transcript

転写付きチャット

プロでロック解除 →

ファイルをここにドラッグまたはクリックしてブラウズ

MP3, WAV, M4A, FLAC, MP4, MKV, MOV, WebM — 最大2GB

複数のファイルを一括アップロードプロと一緒に

アップグレード

Private transcript

転写付きチャット

プロでロック解除 →

アップグレード

リアルタイムの音声からテキストに変換。AI は話すときに自動的に訂正します。長い話をすると正確さが向上します。

まずマイクをテストしてください

10分フリー/日 600分無料クレジットカードなし暗号化

無料登録 →

対応エクスポート形式

音声・動画の文字起こし後、以下の形式でダウンロードできます。すべての形式に完全なテキストが含まれ、タイムド形式にはタイムスタンプが含まれます。

TXT（プレーンテキスト）

.txt

フォーマットなしのシンプルなプレーンテキスト文字起こし。文書、メール、他のアプリケーションへのコピーに最適。話者検出有効時は話者ラベルを含みます。

Free plan

SRT（SubRip字幕）

.srt

最も広くサポートされている字幕形式。連番、タイムスタンプ、テキストを含みます。YouTube、Vimeo、VLC、Premiere Pro、Final Cutなど、ほぼすべてのビデオプレーヤーに対応。

Free plan

VTT（WebVTT）

.vtt

Web Video Text Tracks形式、HTML5ビデオキャプションの標準。スタイリング、ポジショニング、メタデータをサポート。

Basic plan+

DOCX（Word文書）

.docx

見出し、タイムスタンプ、話者ラベル付きのフォーマットされたWord文書。議事録、レポート、Microsoft WordやGoogle Docsでの編集に最適。

Basic plan+

JSON（構造化データ）

.json

単語レベルのタイムスタンプ、信頼度スコア、話者ID、セグメントデータを含む機械可読構造化形式。開発者に最適。

Basic plan+

PDF（ポータブル文書）

.pdf

タイムスタンプ、話者ラベル、STT.aiブランディング付きのプロフェッショナルなPDF。共有、アーカイブ、印刷に最適。

Basic plan+

形式の比較

特徴	TXT	SRT	VTT	DOCX	JSON	PDF
Plain text	✓	✓	✓	✓	✓	✓
Timestamps	✗	✓	✓	✓	✓	✓
Speaker labels	✓	✓	✓	✓	✓	✓
Word-level timing	✗	✗	✗	✗	✓	✗
Confidence scores	✗	✗	✗	✗	✓	✗
Video player compatible	✗	✓	✓	✗	✗	✗
Editable	✓	✓	✓	✓	✓	✗
Machine-readable	✗	✗	✗	✗	✓	✗

どの形式を選ぶべき？

字幕用

Use SRT for maximum compatibility or VTT for web-based video players. SRT works with YouTube, Vimeo, Premiere Pro, Final Cut, and DaVinci Resolve.

文書とレポート用

Use DOCX for editable documents or PDF for sharing and archiving. Both include formatted timestamps and speaker labels.

開発者と統合用

Use JSON for the richest data including word-level timestamps, confidence scores, and speaker IDs. Ideal for building custom applications.

クイックコピペ用

Use TXT for a simple plain text transcript you can paste anywhere -- emails, notes, chat, or any text field.

一括エクスポート

Need to export multiple transcripts at once? STT.ai supports batch export from your transcript library. Select multiple transcripts, choose your format, and download them all in a single ZIP file. Available on all paid plans.

APIエクスポート

Developers can retrieve transcripts in any format via the STT.ai API. Simply specify the desired format in your API request and receive the formatted output directly. The JSON format includes the most detailed data including word-level timestamps and confidence scores.

文字起こしして任意の形式でエクスポート

音声・動画をアップロード。形式を選択。即座にダウンロード。

無料で文字起こしを開始

よくある質問

export formats runs in your browser: paste a URL, upload a file, or record from your mic. STT.ai picks the AI model and returns the transcript in under 5 minutes. Export as TXT, SRT, VTT, DOCX, JSON, or PDF.

Yes — every visitor gets 600 free minutes/month on STT.ai, usable for export formats the same as any other workflow. Paid plans starting at $5/month unlock longer files, private transcripts, and priority queueing.

export formats runs on the same AI models as the rest of STT.ai — our best models reach 95-97% accuracy on clean speech (3-5% Word Error Rate on benchmarks). Switch models on the fly if the first pass is below your target.

export formats can run on any of STT.ai's 10+ models — STT.ai Enhanced (most accurate), Whisper Large V3 (99 languages), NVIDIA Canary (#1 WER on supported langs), Whisper Turbo (fast), Moonshine (lightweight), and more.

Yes. Every transcript exports as SRT or VTT — works with YouTube, Vimeo, TikTok, VLC, and every major video player. The burn-subtitles tool overlays them onto video as hardsubs.

Yes. Speaker diarization automatically labels each voice (Speaker 1, Speaker 2, ...) and you can rename them in the built-in editor. Works across all models and languages.

Most export formats jobs finish in under 5 minutes. A 1-hour audio file typically completes in 2-3 minutes with our fastest models. Speed depends on chosen model and current GPU load.

export formats accepts 20+ formats — MP3, WAV, M4A, FLAC, OGG, MP4, MKV, MOV, WebM, AVI, and more. Output to TXT, SRT, VTT, DOCX, JSON, or PDF.

Yes. Audio files submitted to export formats are processed and deleted by default. Pro plans add client-side encryption — even if STT.ai's database is breached, your transcripts are unreadable without your key. Data is never used for model training without explicit opt-in.

Yes. STT.ai offers a REST API with Python and Node.js SDKs, plus an MCP server for Claude and Cursor — all usable for export formats workflows. Free API tier includes 100 minutes/month.

Yes. Every transcript opens in the built-in editor where you can correct words, rename speakers, adjust timestamps, and add notes. All changes save automatically.

Every transcript gets a unique shareable URL. Export to DOCX or PDF for email. Pro plans add password-protected and permanent links — useful for client work.

STT.ai handles 1,300+ platforms including YouTube, Vimeo, TikTok, SoundCloud, Zoom, Google Meet, podcast hosts, and more. URL transcription works with publicly-available content only — DRM-protected sources can't be transcribed.

エクスポート形式

対応エクスポート形式

TXT（プレーンテキスト）

SRT（SubRip字幕）

VTT（WebVTT）

DOCX（Word文書）

JSON（構造化データ）

PDF（ポータブル文書）

形式の比較

どの形式を選ぶべき？

一括エクスポート

APIエクスポート

文字起こしして任意の形式でエクスポート

よくある質問

How does export formats work on STT.ai?

Is export formats free?

How accurate is export formats?

What AI models can I use for export formats?

Can I get subtitles from export formats?

Does export formats detect different speakers?

How long does export formats take?

What input formats does export formats support?

Is my audio private when I use export formats?

Is there a export formats API?

Can I edit a export formats transcript after?

How do I share what export formats produces?

What other platforms work beyond export formats?