Abstract
Spacecraft engineering documents contain densely coupled entities, abbreviations, and cross-sentence relations that are difficult to convert into knowledge graph triples with rule-based pipelines. This paper presents a controlled large language model (LLM) workflow for triple extraction from long aerospace documents. The method combines sentence-aware chunking under a maximum-length budget, a JSON-constrained prompt, regex-based post-validation, and human-in-the-loop review with entity normalization. Experiments are conducted on a spacecraft corpus containing about 1.2 million Chinese characters from 36 documents and 1,860 paragraphs; 200 paragraphs were manually annotated to form a gold set of 3,842 triples. Compared with a rule/dictionary baseline and a traditional sequence labeling model, the proposed workflow achieves the best precision, recall, and F1, reaching 0.92, 0.87, and 0.89, respectively. An ablation study shows that structured prompting and validation improve parseability from 68% to 99% and raise F1 from 0.78 to 0.89. Sensitivity analysis further indicates that a chunk length of 512 with one-sentence overlap provides the best trade-off between quality and latency in our setting. These results show that controlled LLM extraction is a practical route for building reliable domain knowledge graphs from long engineering documents.
| Original language | English |
|---|---|
| Title of host publication | 2026 International Conference on Complex Systems, Intelligent Science and Computing Technology (CSITec) |
| Editors | Joann Wu |
| Publisher | IEEE |
| Pages | 383-386 |
| Number of pages | 4 |
| ISBN (Print) | 9781971299174 |
| DOIs | |
| Publication status | Published - 27 Mar 2026 |
| Event | 2026 International Conference on Complex Systems, Intelligent Science and Computing Technology (CSITec) - Changsha, Hunan, China Duration: 27 Mar 2026 → 29 Mar 2026 |
Conference
| Conference | 2026 International Conference on Complex Systems, Intelligent Science and Computing Technology (CSITec) |
|---|---|
| Country/Territory | China |
| City | Changsha, Hunan |
| Period | 27/03/26 → 29/03/26 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 9 Industry, Innovation, and Infrastructure
User-Defined Keywords
- spacecraft engineering documents
- large language model
- information extraction
- knowledge graph
- structured output
- json validation
Fingerprint
Dive into the research topics of 'Controlled Llm-Based Triple Extraction from Spacecraft Engineering Documents'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver