Publications-NSC Projects

Article View/Open

Publication Export

Google ScholarTM

NCCU Library

Citation Infomation

Related Publications in TAIR

題名 經由數據取得策略來有效處理產品原生評論中之遺失值
Effective Strategies for Active Feature-Value Acquisition for Managing Missing Values in Online User-Generated Product Reviews
作者 唐揆
貢獻者 企管系
關鍵詞 數據整併和取得; 顧客原生產品評論; 遺失值; 填值; 預測模型
Data mergers and acquisition; User generated product review; Missing values;Imputation; Predictive models
日期 2019-05
上傳時間 4-Jun-2026 13:09:53 (UTC+8)
摘要 在大數據的應用上,有效的數據整併和取得(data mergers and acquisition)是成功的要素,此一 發展引領了一個新興研究領域的發展,其主要探討以改善預測模型準確率為目標的數據取得 策略,本研究計劃主要在延續此一研究領域,探討如何運用數據取得的策略,來處理數據中 遺失值的問題,並將發展出來的方法應用到處理顧客原生產品評論(user generated product reviews, UGPR)中的遺失值。因為一般數據中普遍含有遺失值,如果沒有適當的處理,將產 生在數據分析及推論上的困難及偏差,因而導致決策的錯誤。 在建立預測模型時,數據的遺失值可能會發生在標籤(目標值)或特徵值(預測變數),也可 能兩者都有。主動特徵值取得(active feature-value acquisition, AFA)的研究假設數據的遺失值只 發生在特徵值,並且遺失值的取得需要成本,因此AFA 的目標是找到具有成本效益的數據(遺 失值)取得策略,以達到最高的預測準確度,此一研究議題在金融,市場營銷,品質管制和醫 療等領域均有廣泛的應用。 現有的 AFA 方法多以樣本(instance)為選擇單位,作為遺失值取得的基礎,由於還有許多其他 的可能的選擇方法,這仍是一個很寬廣的的研究領域。在本計劃的第一部分,我們將開發以 特徵值(feature)為選擇基礎的策略,以及設計同時以樣本及特徵值為選擇基礎的策略,此外, 我們也將探討並評估加入填值(imputation)的幾種混合方法。 在計劃的第二部分,我們將AFA 應用在處理UGPR 中的遺失值。最近的調查發現,UGPR 已 被廣泛地運用在電子口碑的廣告行銷、品質管制和產品設計等策略上。UGPR 通常含有大量 的遺失值,例如,在一個產品評論中,整體的產品評價是「非常好」,但卻未對產品的特徵值 作出正面的評價,類似的數據將對建立預測模型產生很大的困難。事實上,這一議題比一般 遺失值的問題更難解決,主要原因是顧客評價之間的相關性,近年來市場學的研究發現產品 評比可能會受到先前評論的影響,例如有研究發現,一些評論者只針對與先前評論有不同看 法的部分發表意見。為了研究這個複雜且重要的問題,並提出有效的AFA 策略,我們將結合 市場學的文獻、AFA 策略和幾個數據挖掘方法,如文字探勘和情緒偵測等,希望對處理UGPR 缺失值的理論和應用能作出實質的貢獻。
Data mergers and acquisition have been commonly used in practice for enhancing big data applications. In response to this development, an emerging research area, namely, active data acquisition (ADA), for improving the performance of predictive models has attracted much attention. The goal of the proposed project is to investigate strategies for handling missing data (values) through data acquisition. It has been well known that missing values often exist in data. Without proper treatments, this could post very serious challenges in data analysis and inference and may result in erroneous decisions. For developing a predictive model, missing values may happen to the class label (target), features (predictors), or both. The research area known as active feature-value acquisition (AFA) considers the scenario, in which missing values happen to features only. It is also assumed that acquiring the true value of a missing value would incur a cost. The goal of AFA is to find cost-effective strategies for acquiring missing data to achieve the maximum prediction accuracy. This problem has applications in the financial, marketing, quality control, and medical areas. The focus of the AFA literature is on instance-based methods, which select instances for acquiring missing values. Since many alternative methods may be considered, this is a rich research area for exploration. We plan to develop selection strategies based on features or a combination of instances and features. Furthermore, we will also explore several hybrid methods of combing AFA and imputation. In the first part of this proposal, we devote our efforts to development and evaluation of such strategies. In the second part, we plan to adapt the methods developed in the first part for treating missing values in reviews shared by product users on the Internet, commonly known as user generated product reviews (UGPR). Recent studies show that UGPR has become an important information source for electronic word-of-mouth (eWOM) advertising, product quality control, and product design. It is evident that UGPR generally contains significant amount of missing values. For example, a review writer may rate a product “very good,” but fail to include any positive comments on the features of the product. This creates a problem in developing a predictive model because of missing values in predictors. In order to address the issue, it is essential to understand the possible causes of missing values in UGPR. Recent studies in the marketing and electronic commerce areas have shown that the product rating of a review may be dependent on prior reviews. For example, it was found that later product ratings tend to differ from previously posted ones. Accordingly, it is easy to find text comments, and their missing values may also depend on prior reviews, as a reviewer may only mention certain features which he or she disagrees on previous reviews. In order to study this very challenging but important problem, it requires connecting the marketing and electronic commerce literatures, AFA strategies, and several data mining methods, such as text mining and sentiment analysis. We hope to contribute to both theory and practice by developing effective methods for treating missing values (information) in UGPR.
關聯 科技部, MOST104-2410-H004-129-MY3, 104.08-107.07
資料類型 report
dc.contributor 企管系
dc.creator (作者) 唐揆
dc.date (日期) 2019-05
dc.date.accessioned 4-Jun-2026 13:09:53 (UTC+8)-
dc.date.available 4-Jun-2026 13:09:53 (UTC+8)-
dc.date.issued (上傳時間) 4-Jun-2026 13:09:53 (UTC+8)-
dc.identifier.uri (URI) https://ah.lib.nccu.edu.tw/item?item_id=182749-
dc.description.abstract (摘要) 在大數據的應用上,有效的數據整併和取得(data mergers and acquisition)是成功的要素,此一 發展引領了一個新興研究領域的發展,其主要探討以改善預測模型準確率為目標的數據取得 策略,本研究計劃主要在延續此一研究領域,探討如何運用數據取得的策略,來處理數據中 遺失值的問題,並將發展出來的方法應用到處理顧客原生產品評論(user generated product reviews, UGPR)中的遺失值。因為一般數據中普遍含有遺失值,如果沒有適當的處理,將產 生在數據分析及推論上的困難及偏差,因而導致決策的錯誤。 在建立預測模型時,數據的遺失值可能會發生在標籤(目標值)或特徵值(預測變數),也可 能兩者都有。主動特徵值取得(active feature-value acquisition, AFA)的研究假設數據的遺失值只 發生在特徵值,並且遺失值的取得需要成本,因此AFA 的目標是找到具有成本效益的數據(遺 失值)取得策略,以達到最高的預測準確度,此一研究議題在金融,市場營銷,品質管制和醫 療等領域均有廣泛的應用。 現有的 AFA 方法多以樣本(instance)為選擇單位,作為遺失值取得的基礎,由於還有許多其他 的可能的選擇方法,這仍是一個很寬廣的的研究領域。在本計劃的第一部分,我們將開發以 特徵值(feature)為選擇基礎的策略,以及設計同時以樣本及特徵值為選擇基礎的策略,此外, 我們也將探討並評估加入填值(imputation)的幾種混合方法。 在計劃的第二部分,我們將AFA 應用在處理UGPR 中的遺失值。最近的調查發現,UGPR 已 被廣泛地運用在電子口碑的廣告行銷、品質管制和產品設計等策略上。UGPR 通常含有大量 的遺失值,例如,在一個產品評論中,整體的產品評價是「非常好」,但卻未對產品的特徵值 作出正面的評價,類似的數據將對建立預測模型產生很大的困難。事實上,這一議題比一般 遺失值的問題更難解決,主要原因是顧客評價之間的相關性,近年來市場學的研究發現產品 評比可能會受到先前評論的影響,例如有研究發現,一些評論者只針對與先前評論有不同看 法的部分發表意見。為了研究這個複雜且重要的問題,並提出有效的AFA 策略,我們將結合 市場學的文獻、AFA 策略和幾個數據挖掘方法,如文字探勘和情緒偵測等,希望對處理UGPR 缺失值的理論和應用能作出實質的貢獻。
dc.description.abstract (摘要) Data mergers and acquisition have been commonly used in practice for enhancing big data applications. In response to this development, an emerging research area, namely, active data acquisition (ADA), for improving the performance of predictive models has attracted much attention. The goal of the proposed project is to investigate strategies for handling missing data (values) through data acquisition. It has been well known that missing values often exist in data. Without proper treatments, this could post very serious challenges in data analysis and inference and may result in erroneous decisions. For developing a predictive model, missing values may happen to the class label (target), features (predictors), or both. The research area known as active feature-value acquisition (AFA) considers the scenario, in which missing values happen to features only. It is also assumed that acquiring the true value of a missing value would incur a cost. The goal of AFA is to find cost-effective strategies for acquiring missing data to achieve the maximum prediction accuracy. This problem has applications in the financial, marketing, quality control, and medical areas. The focus of the AFA literature is on instance-based methods, which select instances for acquiring missing values. Since many alternative methods may be considered, this is a rich research area for exploration. We plan to develop selection strategies based on features or a combination of instances and features. Furthermore, we will also explore several hybrid methods of combing AFA and imputation. In the first part of this proposal, we devote our efforts to development and evaluation of such strategies. In the second part, we plan to adapt the methods developed in the first part for treating missing values in reviews shared by product users on the Internet, commonly known as user generated product reviews (UGPR). Recent studies show that UGPR has become an important information source for electronic word-of-mouth (eWOM) advertising, product quality control, and product design. It is evident that UGPR generally contains significant amount of missing values. For example, a review writer may rate a product “very good,” but fail to include any positive comments on the features of the product. This creates a problem in developing a predictive model because of missing values in predictors. In order to address the issue, it is essential to understand the possible causes of missing values in UGPR. Recent studies in the marketing and electronic commerce areas have shown that the product rating of a review may be dependent on prior reviews. For example, it was found that later product ratings tend to differ from previously posted ones. Accordingly, it is easy to find text comments, and their missing values may also depend on prior reviews, as a reviewer may only mention certain features which he or she disagrees on previous reviews. In order to study this very challenging but important problem, it requires connecting the marketing and electronic commerce literatures, AFA strategies, and several data mining methods, such as text mining and sentiment analysis. We hope to contribute to both theory and practice by developing effective methods for treating missing values (information) in UGPR.
dc.format.extent 116 bytes-
dc.format.mimetype text/html-
dc.relation (關聯) 科技部, MOST104-2410-H004-129-MY3, 104.08-107.07
dc.subject (關鍵詞) 數據整併和取得; 顧客原生產品評論; 遺失值; 填值; 預測模型
dc.subject (關鍵詞) Data mergers and acquisition; User generated product review; Missing values;Imputation; Predictive models
dc.title (題名) 經由數據取得策略來有效處理產品原生評論中之遺失值
dc.title (題名) Effective Strategies for Active Feature-Value Acquisition for Managing Missing Values in Online User-Generated Product Reviews
dc.type (資料類型) report