MrDoc User Guide
Basic concepts
MrDoc page description
2.1 Home page
2.2 Collection browsing page
2.3 Document browsing page
2.4 Document editing and modification
2.4.1 Creating a document
2.4.2 Modify document
2.4.3 Create document template
2.4.4 Modify document template
2.4.5 Brief introduction to Markdown syntax
2.4.7 Uploading attachments
2.4.8 Draw a mind map
2.5, Personal Center
2.5.1 Managing documents
2.5.2 Managing collections
2.5.3 Managing document templates
2.5.4 Managing images
2.5.5 Managing user tokens
2.7 Registration and login
2.7.1 User registration
2.7.2 User login
Getting started
Create collection
Create document
Add member collaboration
Documents and knowledge base basics
Collection basic configuration
Pin collection
WebHook message push
Hide collection on homepage
Collection association set
Collection tab configuration
Document editing and content creation
Add document attachments
Show document directory
Create document shortcut
Set document tags
Set document alias
Insert video
Automatic document saving
Document management and organization
Collection directory sorting
Modify document sorting
Subordinate document control
Set document parent-child
Document version history
Transfer collection
Transfer document
Copy document/Move document
Document access records
Collection export and download
Document download and export
Export document PDF
Document drag-and-drop sorting
Collaboration and permission management
Set document permissions
Set collection access permissions
Document copy protection
Set document watermark
Collection collaboration/Collection member management
Document sharing
Collection sharing
Enable document comments
Image and attachment management
Configure image/attachment upload size limit
Configure image upload formats
Configure attachment upload formats/attachment allowlist
Attachment preview
Transfer attachments/Transfer images
Clean up images
Extract attachment text in bulk with the manage command
Data import and migration
Desktop client import
Import Joplin notebooks
Import Evernote
Web import
Command-line import
Web export
Third-party login configuration
DingTalk QR code login and in-app passwordless login configuration
WeCom authentication integration
LDAP authentication integration configuration
OIDC authentication integration
WeChat Official Account web authorization
Third-party storage configuration
MinIO configuration
Qiniu Cloud OSS configuration
Alibaba Cloud OSS configuration
AWS S3 configuration
AI knowledge base and intelligent Q&A
Basic configuration
AI model configuration
Qdrant deployment
Dify framework configuration
Rebuild document AI index
AI features
AI document creation
AI knowledge base Q&A
AI Q&A bot
WeCom smart robot
OnlyOffice integration
Drawio integration
System settings and management
Site information configuration
Homepage template configuration
Official website theme homepage configuration instructions
User and account configuration
Statistics code configuration
Document ads/info blocks/custom head configuration
Disable update detection
Site-wide search mode
Display image thumbnails in documents
Site feedback
RSS subscription
Site single tag settings
Outbound email configuration
Site data export
Editor configuration
Collection documentation page displays site top navigation bar.
Site announcement configuration
Personal account management
Set default editor
Set user nickname
Modify user password
Bind a third-party account
API and developer interfaces
Get user token
Get collection list
Get collection directory
Get collection document list
Get personal document list
Get specified document content
Create new collection
Create a new document
Update documentation
Upload image
Upload attachments
Verify user token
Upload local Office/PDF files as documents
Client and ecosystem integration
Desktop client
Mobile client
Browser extensions
Obsidian plugins
Common usage issues index
Powered by MrDoc Pro
-
+
home
Extract attachment text in bulk with the manage command
# Batch Extract Attachment Text Content Using the manage Command Command: `extract_attachment_content` ## 1. Command Introduction Populates searchable text content for existing historical attachments on the site, for use in attachment preview, site search, and AI knowledge base Q&A. | Item | Description | | --- | --- | | Target | Sites that have been upgraded to the new version and contain historical attachments | | Number of runs | Once is sufficient; already-processed attachments are skipped automatically | | Execution location | MrDoc installation directory | | Execution time | Off-peak hours | | Service status | Keep running; no need to stop the service | | Data safety | A database backup is recommended beforehand | Constraints: - Supported types for text extraction: pdf, doc, docx, xls, xlsx, pptx; other types are skipped. - `.doc` depends on LibreOffice on the server (converted before reading): the Docker image includes it; source deployments must install it themselves, and specify the path to its executable in the `preview.libreoffice_path` setting in the `config.ini` configuration file (default `soffice`). If not installed, `.doc` will be extracted as empty; `.docx` is unaffected. - Attachments on cloud storage (Qiniu, MinIO, etc.) are not processed. - Overly long content is truncated at the limit; the complete content is subject to the original attachment. - Already-processed attachments are not read again; add `--force` to re-read them. - The AI index must be rebuilt before AI knowledge base Q&A will read the attachment content. - Do not interrupt the process while the command is running. ## 2. Usage Syntax ```sh python manage.py extract_attachment_content [参数] ``` Docker deployment (the container name is subject to what `docker ps` actually shows): ```sh docker exec -it mrdoc python manage.py extract_attachment_content [参数] ``` ## 3. Parameter Description | Parameter | Description | | --- | --- | | No parameter | Process all attachments without content | | `--dry-run` | Inventory only, no writes | | `--limit N` | Process at most N attachments | | `--force` | Re-read attachments that already have content | | `--attachment-id 1,2,3` | Process only the specified attachments, comma-separated | | `--doc-id N` | Process only attachments referenced by the specified document | | `--reindex` | Rebuild the AI index of affected documents after completion | ## 4. Usage Examples ### Example 1: Inventory attachments to be processed ```sh python manage.py extract_attachment_content --dry-run ``` Check the number of "attachments to be processed" in the output, and continue once it matches expectations. ### Example 2: Full completion (AI Q&A not enabled) ```sh python manage.py extract_attachment_content ``` ### Example 3: Full completion and AI index rebuild (AI Q&A enabled) ```sh python manage.py extract_attachment_content --reindex ``` ### Example 4: Targeted processing of attachments for a single document ```sh python manage.py extract_attachment_content --doc-id 43090 ``` ### Output Reference ``` 待处理附件:70 个 提取成功:12 | attachment/2025/03/xxx.pdf | 长度=12345 跳过(不支持格式):13 | xxx.zip 跳过(非本地存储/文件缺失):14 | http://... ========== 提取完成 ========== 处理附件: 70 提取成功: 67 仍为空(提取失败): 1 跳过-不支持的格式: 2 跳过-非本地存储/文件缺失: 0 ``` Additional output when using `--reindex`: ``` 开始重建AI索引,受影响文档 5 篇 索引重建成功:DocID=43090 | xxx AI索引重建完成:成功=5 失败=0 ``` | Output item | Action | | --- | --- | | Extraction succeeded | No action needed | | Still empty (extraction failed) | No action needed (scanned PDFs, encrypted or corrupted files) | | Skipped - unsupported format | No action needed | | Skipped - non-local storage/file missing | Record the count and contact technical support to confirm the attachment storage configuration | | Message | Meaning | | --- | --- | | No attachments to be processed | No attachments to be completed | | AI index scope is an empty whitelist | The knowledge scope is not set in the backend "AI Configuration"; it must be set first | | No published documents referencing these attachments found | No need to rebuild the index | ### Verification 1. Open a document containing attachments and check whether the attachment preview displays the body text; 2. Search for a passage of text within the attachment in site search and confirm it can be found; 3. On sites with AI Q&A enabled, search for the attachment content in AI Q&A and confirm it can be found. ### Reporting Anomalies If the command errors out and is interrupted or produces output that cannot be interpreted, record the following information and contact technical support: 1. The complete error output; 2. The execution time and the parameters used; 3. The information of the last attachment being processed.
mrdoc
Sept. 29, 2026, 6:41 p.m.
Forward
Favorites
Last
Next
Scan the QR Code
Copy link
Scan the QR code to share.
Copy link
share
link
type
password
Update password
Validity period
Markdown file
Word document
PDF document (print)