文本的数据分析原则

政务大数据应用与分析 (80700673)

胡悦

清华大学社会科学学院

个人简介

个人经历

  • 政治学博士(University of Iowa)
    • 信息学(Graduated Certificate in Informatics)
  • 清华大学计算社会科学平台(副主任)
    • 清华数据与治理中心(副主任)
    • 计算社会科学编程语言证书项目(负责人)
    • Learning R with Dr. Hu & Friends 工作坊(创始人)

研究兴趣:认知、行为与现代性

  • 方法路径:计算政治学
    • 实验室和调查实验
    • 潜变量分析、网络分析、空间分析
    • 文本大数据分析、数据可视化

研究领域:比较政治、国家治理

  • W. 心理学
    • 政治认知治理
    • 行为公共政策
    • 政治传播
  • W. 经济学
    • 经济不平等感知
    • 公共设施、服务均等化
  • W. 语言学
    • 权力背书的语言效果与机制
    • 语言政策的治理功能

提要

King (2015): [The big-data approach is] the end of the quantitative-qualitative divide.

认知

见字如数
文本与文本数据

  • 文本→数据
  • 文本分析
  • 应用案例

原则

以数知文:
文本数据化原理规范

  • 数据获取
  • 数据整理
  • 基本分析

操作

望数生义:
文本数据分析过程示范(R)

见字如数:建立认知

文本数据

每年: 2024年初社交媒体用户50.4亿,2023年新增2.66亿首次使用用户 (Kemp 2024)

每日:

  • 百度日用户搜索请求,需1.7天才能扫描一遍;
  • 微信日增数据500TB——比人类所有书籍存量还多。

每秒:全世界每秒发送290万封email,一人需要5.5年日以继夜才能读完。

Tip

2023年,全国数据生产总量达32.85ZB,,同比增长22.44%,即人均 22 TB (颜之宏 and 严赋憬 2024)1

预计2027年,全球非结构化数据将占到数据总量的86.8% (IDC FutureScape 2024)

文本研究

历史悠久而非主流 ← 资料难获取;花时间;难推广;难管理;难分析

新兴工具的繁荣:

  • 文本资料指数级增长;
  • 大规模文本数据采集
  • 存储和管理能力增强;
  • 可推广、系统化和廉价化;
  • 文本分析方法蓬勃发展

(计算机辅助)文本分析

对象

文字 语言

类型

  • Text analysis vs. content analysis
  • Representational analysis vs. Instrumental analysis
  • Thematic analysis vs. semantic analysis

🌰 I

Grimmer (2010)

  • Objective
    • The priorities political actors emphasize in statements
  • Data
    • An original collection of over 24,000 Senate press releases in 2007
  • Method
    • Bayesian Hierarchical Topic Model

发现

Committee leaders focus on their committee’s issues

Committee leaders focus on their committee’s issues

Expressed agendas cluster geographically

Expressed agendas cluster geographically

Attention to appropriations predicts opposition to earmark reform

Attention to appropriations predicts opposition to earmark reform

🌰 II

Benoit et al. (2016)

  • Objective
    • Professionals vs. Masses
  • Method
    • Crowd-sourced identification
  • Data
    • 18,263 natural sentences from British Conservative, Labour and Liberal Democrat manifestos

操作

🌰 III

Dietrich, Hayes, and O’Brien (2019)

  • Objective
    • Speakers’ emotional state
  • Method
    • Analyses of vocal pitch
  • Data
    • 74,158 Congressional floor speeches

操作

🌰 IV

Zhang and Pan (2019)

  • Objective
    • Group activities from social media
  • Method
    • CNN for images; CNN-RNN for texts
  • Data
    • A random sample of 20,000 geocoded posts from Weibo, 2010–2017

操作

挑战

理论: Single causal mechanism

前提: Bag of words

数据

  • DGP
    • 随机抽取
    • Only posted
    • One-time censor
  • 非结构化
  • 海量潜在维度
  • 内容复杂且微妙

以数知文: 理解原则

本节要回答之问题

  1. 资源哪里找?
  2. 信息如何用?
  3. 数据能干啥?
  4. 红线在哪里

原生数据

  • Email/短信
  • 网站HTML
  • RSS feeds
  • 网络社交媒体
  • 网络论坛
  • 网络问答平台
  • 媒体数据库
  • 网络交易行为
    ……

社交媒体

社交媒体

公共开放平台

网络问政平台

社会化问答网站

媒体数据库

问卷开放性问题

二手数据

  • 中国知网等数据库(期刊、报刊、年鉴等)
  • Google Books、百度学术
  • Google Trend、百度指数
  • JSTOR Data for Research……

文本获取

  • 原生数据:Spider/crawler/scraper
  • 二手数据:档案数据和数字化数据

操作演示

编程抓取

SelectorGadget (Chrome add-in)

Scrapping with rvest

ls_zhongsheng <-
  read_html("http://politics.people.com.cn/GB/8198/426918/index.html") |> # index page
  html_nodes("h5 a") |> # the nodes of the links
  html_attr("href") |> # just the links
  str_replace("^/n1", "http://politics.people.com.cn/n1")

df_zhongsheng <- map_df(ls_zhongsheng, function(link) {
  zs_article <- read_html(link, encoding = "GB18030") # read the html
  
  zs_title <- html_nodes(zs_article, "h1") |>
    html_text
  
  zs_time <- html_nodes(zs_article, ".box01 .fl") |>
    html_text |>
    str_extract("\\d{4}年\\d{2}月\\d{2}日")
  
  zs_content <- html_nodes(zs_article, "#rwb_zw p") |>
    html_text |>
    str_c(collapse = "") |> # combined the paragraphs
    str_remove_all("\\s|\\n|\\t") # remove the horizontal spaces
  
  zs_data <- data.frame(title = zs_title,
                        time = zs_time,
                        content = zs_content)
})

正则表达式

文本数据结构化

文本分析的基础原则 (Grimmer and Stewart 2013)

  1. All quantitative models of language are wrong—but some are useful.
  2. Quantitative methods for text amplify resources and augment humans.
  3. There is no globally best method for automated text analysis.
  4. Validate, Validate, Validate.

望数生义

文本分析(传统)方法概览

Grimmer and Stewart (2013)

Grimmer and Stewart (2013)

研究层次

数据层次

  • Corpus
    • Document (volumn, chapter, section)
      • Paragraph
        • Sentence
          • Clause
            • Word (Unigram)
              • Token

“Token”:语言特征单元

  • Token in a document: term
  • Token in a group: N-gram, e.g., a word = a Unigram

分析层次

描述

词频、词云(👎)、网络

聚类

知类分文、知文分类

语义

情感分析(sentiment analysis)

A Bag of Words

All quantitative models of language are wrong—but some are useful (Grimmer and Stewart 2013).

Bag of words (BoW)

A text is represented as the bag (multiset) of its words.

从自然语言到计算机语言

Document-Term Matrix (DTM)

Document-Term Matrix (DTM)

预处理

  • Segmentation
  • Tokenization
  • Stopwords (停词)/function words removing
  • 其他根据研究目的的删减

Segmentation

Scriptio discreta (e.g., English)

Document → paragraphs/sentences

Scriptio continua (e.g, CJK)

 [1] "近年来"   "美国"     "一些"     "政客"     "被"       "美国"     "优先"    
 [8] "遮住"     "了"       "双眼"     "大"       "搞"       "贸易"     "保护主义"
[15] "单边主义" "肆意"     "挥舞"     "关税"     "大棒"     "全然不顾" "中美"    
[22] "两国人民" "和"       "全世界"   "人民"     "的"       "强烈"     "反对"    

Tokenization

  • 目标:去除Syntax
    • 大小写
    • 标点
    • 非字符(@#¥%……&*)
    • 停词
 [1] "to"     "can"    "could"  "dare"   "do"     "did"    "does"   "may"    "might" 
[10] "would"  "should" "must"   "will"   "ought"  "shall"  "need"   "is"     "a"     
[19] "am"     "are"    "about" 
 [1] "一些"   "一何"   "一切"   "一则"   "一方面" "一旦"   "一来"   "一样"   "一般"  
[10] "一转眼" "万一"   "上"     "上下"   "下"     "不"     "不仅"   "不但"   "不光"  
[19] "不单"   "不只"   "不外乎"

中文停词表

Scriptio discreta special

  • Lemmatization: (happy, happier, happiness) → happy
  • Stemming: (happy, happier, happiness) → happi

More Examples

  • Lemmization: (went, leaves, geese, unhappy) → (go, leaf, goose, unhappy)
  • Stemming: (went, leaves, geese, unhappy) → (went, leav, geese , unhappi)

为什么要这么做?

还能做什么

  • Labeling
    • Content vs. function
    • Linguistic features: n., v., adj., adv., prep., conj…….

综合的🌰

2019-05-14 ~ 05-22: 中美贸易战

2019-05-14 ~ 05-22: 中美贸易战

数据清理

原始数据

                                           title       time word
1 难道非要撞了南墙才回头——意孤行必将失败(钟声) 2019-05-30   一
2 难道非要撞了南墙才回头——意孤行必将失败(钟声) 2019-05-30   场
3 难道非要撞了南墙才回头——意孤行必将失败(钟声) 2019-05-30 经贸
4 难道非要撞了南墙才回头——意孤行必将失败(钟声) 2019-05-30 摩擦
5 难道非要撞了南墙才回头——意孤行必将失败(钟声) 2019-05-30   让
6 难道非要撞了南墙才回头——意孤行必将失败(钟声) 2019-05-30 世界

去掉停词

                                           title       time word
1 难道非要撞了南墙才回头——意孤行必将失败(钟声) 2019-05-30   场
2 难道非要撞了南墙才回头——意孤行必将失败(钟声) 2019-05-30 经贸
3 难道非要撞了南墙才回头——意孤行必将失败(钟声) 2019-05-30 摩擦
4 难道非要撞了南墙才回头——意孤行必将失败(钟声) 2019-05-30 世界
5 难道非要撞了南墙才回头——意孤行必将失败(钟声) 2019-05-30 看到
6 难道非要撞了南墙才回头——意孤行必将失败(钟声) 2019-05-30 一个

词频分析

词频异质性

相关性分析

关键词识别与分析

2019-05-07 关键词

$simhash
[1] "4937328317770165287"

$keyword
 101.845  62.3717   58.696  57.5996  53.6127 
  "增长"   "韧性"    "GDP" "一季度"   "经济" 

2019-05-30 关键词

$simhash
[1] "9551059633200593385"

$keyword
99.7583 56.3389 37.4295 35.7857 34.9402 
 "美国"  "政客"  "别国"  "国际"  "世界" 

相似度分析

通过“距离”测量相似度

Hamming distance: the distance between two strings of equal length is the number of positions at which the corresponding symbols are different.

e.g., H(100→011) = 3; H(010→111) = 2.

相似度能告诉你什么

05-30 vs. 05-22

$distance
[1] 30

$lhs
99.7583 56.3389 37.4295 35.7857 34.9402 
 "美国"  "政客"  "别国"  "国际"  "世界" 

$rhs
171.456 135.966 123.366 114.446 93.1847 
 "磋商"  "美方"  "中方"  "背弃"  "倒退" 

05-30 vs. 05-23

$distance
[1] 18

$lhs
99.7583 56.3389 37.4295 35.7857 34.9402 
 "美国"  "政客"  "别国"  "国际"  "世界" 

$rhs
   137.594    91.4451    63.4508    53.6785    37.6035 
    "规则"     "美国"     "美方"     "国际" "世贸组织" 

主题模型(Topic modeling)

总结

  1. 认知
    • 丰富资源
    • 技术门槛
  2. 原则
    • 在“错误”的前提下寻找价值
  1. 操作
    • 打散:预处理与结构化
    • 聚合:
      • 词频
      • 相关性/相似度
      • 主题模型

Distant reading

Distant reading

感谢倾听,欢迎交流

  sammo3182

  yuehu@tsinghua.edu.cn

  https://www.drhuyue.site

参考文献

Benoit, Kenneth, Drew Conway, Benjamin E. Lauderdale, Michael Laver, and Slava Mikhaylov. 2016. “Crowd-Sourced Text Analysis: Reproducible and Agile Production of Political Data.” American Political Science Review 110 (2): 278–95. https://doi.org/10.1017/S0003055416000058.
Coase, R. H. 1960. “The Problem of Social Cost.” The Journal of Law & Economics 56 (4): 837–77. https://doi.org/10.1086/674872.
Dietrich, Bryce J., Matthew Hayes, and Diana Z. O’Brien. 2019. “Pitch Perfect: Vocal Pitch and the Emotional Intensity of Congressional Speech.” American Political Science Review, Forthcoming. http://www.brycejdietrich.com/files/working_papers/DietrichHayesOBrien.pdf.
Friedman, Milton, and Anna Jacobson Schwartz. 2008. A Monetary History of the United States, 1867-1960. Princeton University Press. https://books.google.com?id=Q7J_EUM3RfoC.
Grimmer, Justin. 2010. “A Bayesian Hierarchical Topic Model for Political Texts: Measuring Expressed Agendas in Senate Press Releases.” Political Analysis 18 (1): 1–35. https://doi.org/10.1093/pan/mpp034.
Grimmer, Justin, and Brandon M. Stewart. 2013. “Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts.” Political Analysis 21 (3): 267–97.
IDC FutureScape. 2024. “2024年中国数据和分析市场十大预测.” IDC Media Center. https://www.idc.com/getdoc.jsp?containerId=prCHC51814824.
Kemp, Simon. 2024. “Digital 2024: Global Overview Report.” Meltwater. https://runwise.oss-accelerate.aliyuncs.com/sites/15/2024/04/2024%E5%B9%B4%E5%85%A8%E7%90%83%E6%95%B0%E5%AD%97%E5%8C%96%E8%90%A5%E9%94%80%E6%B4%9E%E5%AF%9F%E6%8A%A5%E5%91%8A-50%E4%BA%BF%E7%A4%BE%E4%BA%A4%E5%AA%92%E4%BD%93%E7%94%A8%E6%88%B7_Meltwater%E8%9E%8D%E6%96%87_2024.pdf.
King, Gary. 2015. “Big Data Is Not about the Data!” Guest talk presented at the Talk at the capital markets cooperative research centre, Sydney, Australia, November 11. https://gking.harvard.edu/files/gking/files/evbase-cmcrc.pdf.
Zhang, Han, and Jennifer Pan. 2019. “CASM: A Deep-Learning Approach for Identifying Collective Action Events with Text and Image Data from Social Media.” Sociological Methodology 49 (1): 1–57. https://doi.org/10.1177/0081175019860244.
颜之宏, and 严赋憬. 2024. “最新报告出炉!2023年我国数据生产总量达32.85ZB.” 新华网. May 24, 2024. http://www.news.cn/20240524/a1b191e10fa84fa0af52ba88b6dcb1b7/c.html.