Skip to main content

Research Repository

Advanced Search

CDTier:A Chinese Dataset of Threat Intelligence Entity Relationships

Zhou, Yinghai; Ren, Yitong; Yi, Ming; Xiao, Yanjun; Tan, Zhiyuan; Moustafa, Nour; Tian, Zhihong


Yinghai Zhou

Yitong Ren

Ming Yi

Yanjun Xiao

Nour Moustafa

Zhihong Tian


Cyber Threat Intelligence (CTI), which is knowledge of cyberspace threats gathered from security data, is critical in defending against cyberattacks.However, there is no open-source CTI dataset for security researchers to effectively apply enormous CTI information for security analysis in the field of threat intelligence, particularly in the field of Chinese threat intelligence. As a result, for network security research and development, this paper constructed a Chinese CTI entity relationship dataset–CDTier, which includes: 1) A threat entity extraction dataset composed of 100 CTI reports, 3744 threat sentences and 4259 threat knowledge objects; 2) A dataset for entity relation extraction including 100 CTI reports, 2598 threat sentences and 2562 knowledge object relations. CDTier is, as far as we know, the first CTI dataset. On the CDTier, we trained 4 models for threat entity extraction and relation extraction using well-established and widely used deep learning methods in the NLP. The results showed that the model trained on CDTier extracts knowledge objects and their relationships described in threat intelligence more accurately. This significantly minimizes threat intelligence analysts' work while assessing threat intelligence. The CDTier may be found at .

Journal Article Type Article
Acceptance Date Jan 23, 2023
Online Publication Date Jan 30, 2023
Publication Date 2023
Deposit Date Jan 31, 2023
Publicly Available Date Jan 31, 2023
Electronic ISSN 2377-3782
Publisher Institute of Electrical and Electronics Engineers
Peer Reviewed Peer Reviewed
Volume 8
Issue 4
Pages 627-638
Keywords Cyber threat intelligence, entity relation extraction, information extraction, NLP, threat entity extraction
Public URL


CDTier: A Chinese Dataset Of Threat Intelligence Entity Relationships (accepted version) (12.1 Mb)

You might also like

Downloadable Citations