Assessing Effectiveness of Using Internal Signals for Check-Worthy Claim\n Identification in Unlabeled Data for Automated Fact-Checking
作者:Archita Pathak, Rohini K. Srihari · 发表于:arXiv (Cornell University) · 年份:2021 · DOI:10.48550/arxiv.2111.01706 · 被引用次数:1 · 研究领域:Topic Modeling、Explainable Artificial Intelligence (XAI)、Software Engineering Research
While recent work on automated fact-checking has focused mainly on verifying\nand explaining claims, for which the list of claims is readily available,\nidentifying check-worthy claim sentences from a text remains challenging.\nCurrent claim identification models rely on manual annotations for each\nsentence in the text, which is an expensive task and challenging to conduct on\na frequent basis across multiple domains. This paper explores methodology to\nidentify check-worthy claim sentences from fake news articles, irrespective of\ndomain, without explicit sentence-level annotations. We leverage two internal\nsupervisory signals - headline and the abstractive summary - to rank the\nsentences based on semantic similarity. We hypothesize that this ranking\ndirectly correlates to the check-worthiness of the sentences. To assess the\neffectiveness of this hypothesis, we build pipelines that leverage the ranking\nof sentences based on either the headline or the abstractive summary. The\ntop-ranked sentences are used for the downstream fact-checking tasks of\nevidence retrieval and the article's veracity prediction by the pipeline. Our\nfindings suggest that the top 3 ranked sentences contain enough information for\nevidence-based fact-checking of a fake news article. We also show that while\nthe headline has more gisting similarity with how a fact-checking website\nwrites a claim, the summary-based pipeline is the most promising for an\nend-to-end fact-checking system.\n