Publications-NSC Projects

Article View/Open

Publication Export

Google ScholarTM

NCCU Library

Citation Infomation

Related Publications in TAIR

題名 中文讀者真實性判斷研究及其應用
An investigation of Chinese readers' veridicality judgments and its applications
作者 張瑜芸
貢獻者 語言所
關鍵詞 事件; 真實性判斷; 語境; 語言特徵; 讀者; 語料庫
event; veridicality; context; linguistic features; readers; corpus
日期 2022-01
上傳時間 23-Jul-2026 15:24:04 (UTC+8)
摘要 本計畫從讀者角度出發,以語言學角度去探索新聞文本的語言特徵之間的互動性如何影響讀者相信該事件會發生 (reader's veridicality judgment),並看這些語言特徵是否能幫助機器學習模型自動分類及預測讀者對於事件真實性之判斷。在這瞬變萬化的網路世界裡,能更自動分類事件真實性判斷是一件重要工程。然而,目前大多資訊抽取系統只能擷取子句層次的訊息,但這樣的訊息對於讀者來說是不足以代表他們對於整個事件的理解程度,例如這兩個句子中 ``The FBI alleged in court documents that Zazi had admitted having a handwritten recipe for explosives on his computer'' 和 ``According to the FBI agents, there is relatively little evidence that Zazi had a handwritten recipe for explosives'' 皆只能抽取出子句 ``Zazi had a handwritten recipe for explosives'' 訊息。先前研究並沒有針對語言特徵之間的互動性進行勘查,因此此計畫將透過拆解讀者和新聞事件之日常互動關係,將更深一層隱含訊息增加至目前事件抽取系統。計畫中將採用兩個機器學習模型,Maximum Entropy (MaxEnt) 和 Long Short-Term Memory (LSTM)。本計畫也提出新的量表,除了包含邏輯及語言學上的考量以外,也將實際應用層面納入考量。此外,研究也將釋出一個開源工具,以去年剛釋出的CKIP CoreNLP 為基礎開發採語言學規則建立之事件抽取系統。另外,以半監督式方式結合最新的聚集分類 (clustering) 方式擴充的語料庫也將開放免費使用。由於現階段深度學習模型除了往語義、句法、甚至隱含訊息的語境學習訓練前進,但供其測試用且含有語境訊息標記資料集少,尤其是中文,故此擴充後的語料庫能作為一份資源,提供深度學習模型做更深層的認知知識訓練及測試用。
The goal of this project is to explore how the interactions of pragmatically enriched linguistic cues can assist machine learning models to classify and predict readers' veridicality judgments automatically, and how the collected corpus (the Chinese PragBank) can be further adapted into real world applications. Readers' veridicality judgments are whether readers view an event described in a sentence as happening or not. For instance, in ``The FBI alleged in court documents that Zazi had admitted having a handwritten recipe for explosives on his computer'', do people believe that Zazi had a handwritten recipe for explosives? On the other hand, what do people infer if the sentence is ``According to the FBI agents, there is relatively little evidence that Zazi had a handwritten recipe for explosives''? Automatically classifying veridicality of events is important to swift through the ever growing amount of information appearing online. However, most information extraction systems nowadays work roughly at the clause level, and would extract that ``Zazi had a handwritten recipe for explosives'' in both sentences given above. However, for readers, the above information is not enough for representing the whole picture of an event while reading daily news. Previous studies related to event extraction did not consider the interplay of an event and other factors which affects readers' veridicality judgments. Through investigating the relationships among event-related features, this project reveals the daily interactions of news events and readers, which will enrich the current event extraction systems with deeper implicit information. Two machine learning models, Maximum Entropy (MaxEnt) and Long Short-Term Memory (LSTM), will be implemented in this project. A 5-interval veridicality scale is proposed in the project, which not only takes logical and linguistic perspectives into account, but also its practical usage in real world implementation. The scale will serve as a benchmark for researchers who work on modeling this kind of pragmatic classification task. A rule-based event extraction toolkit will be open-sourced for extracting event-related information from Chinese news texts based on parsed results produced by CKIP CoreNLP. To the best of my knowledge, there is no such toolkit released yet on top of the outputs from CKIP CoreNLP, which is just newly open-sourced last year. In addition, an expanded Chinese PragBank corpus will be open-sourced for researchers to access as well, created via the proposed semi-supervised approach proposed in this project with the state-of-the-art clustering model. With the rise of deep learning (DL) models, more tasks challenge on pushing models to learn semantic, syntactic, and even implicit information embedded within context. However, not a lot of semantically and pragmatically evaluated datasets are open-accessed, particularly in Chinese. Thus, this expanded corpus will serve as a resource for DL models to train as well as evaluate on deeper level information.
關聯 科技部, MOST109-2410-H004-191, 109.11-110.10
資料類型 report
dc.contributor 語言所
dc.creator (作者) 張瑜芸
dc.date (日期) 2022-01
dc.date.accessioned 23-Jul-2026 15:24:04 (UTC+8)-
dc.date.available 23-Jul-2026 15:24:04 (UTC+8)-
dc.date.issued (上傳時間) 23-Jul-2026 15:24:04 (UTC+8)-
dc.identifier.uri (URI) https://ah.lib.nccu.edu.tw/item?item_id=183426-
dc.description.abstract (摘要) 本計畫從讀者角度出發,以語言學角度去探索新聞文本的語言特徵之間的互動性如何影響讀者相信該事件會發生 (reader's veridicality judgment),並看這些語言特徵是否能幫助機器學習模型自動分類及預測讀者對於事件真實性之判斷。在這瞬變萬化的網路世界裡,能更自動分類事件真實性判斷是一件重要工程。然而,目前大多資訊抽取系統只能擷取子句層次的訊息,但這樣的訊息對於讀者來說是不足以代表他們對於整個事件的理解程度,例如這兩個句子中 ``The FBI alleged in court documents that Zazi had admitted having a handwritten recipe for explosives on his computer'' 和 ``According to the FBI agents, there is relatively little evidence that Zazi had a handwritten recipe for explosives'' 皆只能抽取出子句 ``Zazi had a handwritten recipe for explosives'' 訊息。先前研究並沒有針對語言特徵之間的互動性進行勘查,因此此計畫將透過拆解讀者和新聞事件之日常互動關係,將更深一層隱含訊息增加至目前事件抽取系統。計畫中將採用兩個機器學習模型,Maximum Entropy (MaxEnt) 和 Long Short-Term Memory (LSTM)。本計畫也提出新的量表,除了包含邏輯及語言學上的考量以外,也將實際應用層面納入考量。此外,研究也將釋出一個開源工具,以去年剛釋出的CKIP CoreNLP 為基礎開發採語言學規則建立之事件抽取系統。另外,以半監督式方式結合最新的聚集分類 (clustering) 方式擴充的語料庫也將開放免費使用。由於現階段深度學習模型除了往語義、句法、甚至隱含訊息的語境學習訓練前進,但供其測試用且含有語境訊息標記資料集少,尤其是中文,故此擴充後的語料庫能作為一份資源,提供深度學習模型做更深層的認知知識訓練及測試用。
dc.description.abstract (摘要) The goal of this project is to explore how the interactions of pragmatically enriched linguistic cues can assist machine learning models to classify and predict readers' veridicality judgments automatically, and how the collected corpus (the Chinese PragBank) can be further adapted into real world applications. Readers' veridicality judgments are whether readers view an event described in a sentence as happening or not. For instance, in ``The FBI alleged in court documents that Zazi had admitted having a handwritten recipe for explosives on his computer'', do people believe that Zazi had a handwritten recipe for explosives? On the other hand, what do people infer if the sentence is ``According to the FBI agents, there is relatively little evidence that Zazi had a handwritten recipe for explosives''? Automatically classifying veridicality of events is important to swift through the ever growing amount of information appearing online. However, most information extraction systems nowadays work roughly at the clause level, and would extract that ``Zazi had a handwritten recipe for explosives'' in both sentences given above. However, for readers, the above information is not enough for representing the whole picture of an event while reading daily news. Previous studies related to event extraction did not consider the interplay of an event and other factors which affects readers' veridicality judgments. Through investigating the relationships among event-related features, this project reveals the daily interactions of news events and readers, which will enrich the current event extraction systems with deeper implicit information. Two machine learning models, Maximum Entropy (MaxEnt) and Long Short-Term Memory (LSTM), will be implemented in this project. A 5-interval veridicality scale is proposed in the project, which not only takes logical and linguistic perspectives into account, but also its practical usage in real world implementation. The scale will serve as a benchmark for researchers who work on modeling this kind of pragmatic classification task. A rule-based event extraction toolkit will be open-sourced for extracting event-related information from Chinese news texts based on parsed results produced by CKIP CoreNLP. To the best of my knowledge, there is no such toolkit released yet on top of the outputs from CKIP CoreNLP, which is just newly open-sourced last year. In addition, an expanded Chinese PragBank corpus will be open-sourced for researchers to access as well, created via the proposed semi-supervised approach proposed in this project with the state-of-the-art clustering model. With the rise of deep learning (DL) models, more tasks challenge on pushing models to learn semantic, syntactic, and even implicit information embedded within context. However, not a lot of semantically and pragmatically evaluated datasets are open-accessed, particularly in Chinese. Thus, this expanded corpus will serve as a resource for DL models to train as well as evaluate on deeper level information.
dc.format.extent 116 bytes-
dc.format.mimetype text/html-
dc.relation (關聯) 科技部, MOST109-2410-H004-191, 109.11-110.10
dc.subject (關鍵詞) 事件; 真實性判斷; 語境; 語言特徵; 讀者; 語料庫
dc.subject (關鍵詞) event; veridicality; context; linguistic features; readers; corpus
dc.title (題名) 中文讀者真實性判斷研究及其應用
dc.title (題名) An investigation of Chinese readers' veridicality judgments and its applications
dc.type (資料類型) report