| A | B | C | D | E | |
|---|---|---|---|---|---|
| Big | 30 | 13 | 1 | 16 | 23 |
| Medium | 180 | 121 | 111 | 187 | 138 |
| Small | 97 | 593 | 862 | 20 | 4 |
计算政治学讲习班
清华大学社会科学学院
Vaild [latent variable] measurement is the cornerstone of successful scientific inquiry (Carpini and Keeter 1993, 1203).
指标问题(1~10):
累加综合法(additive scales)
\[\tilde{X} = (X_1 + X_2 + X_3)/3.\]
潜在问题
如何做得更好?
连续因子模型
离散回应模型
待解之题
\[Y_{kq} = \frac{\sum Y_{ikq}}{n}.\]
不妥之处?
\[ \theta_h = \frac{\sum_{j \in h} N_j \color{red}{\mu_j} }{\sum_{j \in h} n_j}, \]
N: 总体(来自普查)
n: 样本(来自sample)
数据:某年某市五区域2396家产业公司的财政信息
目标:估测每个区域的产业平均收入
公司规模和区域分布
| A | B | C | D | E | |
|---|---|---|---|---|---|
| Big | 30 | 13 | 1 | 16 | 23 |
| Medium | 180 | 121 | 111 | 187 | 138 |
| Small | 97 | 593 | 862 | 20 | 4 |
总体平均值(真值)
| Zone | income |
|---|---|
| A | 652.28 |
| B | 320.75 |
| C | 331.02 |
| D | 684.98 |
| E | 767.39 |
我们随机选取数据中1000个产业公司作为样本。 如果直接各组取平均值的话就会产生偏差:
Step I: Mr \[ \begin{align} 收入 = \beta_{0z}& + \beta_{1Z=z}公司规模_{iz} + \epsilon_{iz}, \\ \beta_{0z}& = \gamma_{00} + \gamma_{01}区域_z + u_{0z}. \end{align} \]
Output: Post-strata means (\(\mu_z\))
| A | B | C | D | E | |
|---|---|---|---|---|---|
| Big | 1274.74 | 1148.58 | 1189.59 | 1238.51 | 1251.95 |
| Medium | 706.19 | 580.03 | 621.03 | 669.96 | 683.40 |
| Small | 372.95 | 246.79 | 287.79 | 336.72 | 350.16 |
Step II: P
A B C D E
656.4551 318.3753 326.6951 680.8675 754.5761
MrP 没有解决的问题
→ 群组IRT
Conquer the individual fallacy
Dynamic Group-level IRT(DGIRT)
群组(Group-level):1
\(\eta_{ktq} = logit^{-1}(\frac{\color{red}{\bar{\theta}_{kt}}- \beta_q}{\sqrt{\alpha^2_q + \color{red}{(1.7\sigma_{kt})^2}}}),\)
\(\bar{\theta}_k\) 和 σkt 是潜在变量在组k时间t的均值和SD。
纳入时空 (Dynamic)
\[ \begin{align} \bar{\theta}_k\sim& N(\xi_t + \boldsymbol{x'_k\gamma}, \sigma^2_{\bar{\theta}}), \\ \xi_t \sim& N(\xi_{t-1}, \sigma^2_{\gamma}), \\ \gamma_{pt} \sim& N(\gamma_{p,t-1}, \delta_t + \boldsymbol{z'_p.\eta_t}, \sigma^2_{\gamma}) \end{align} \]
效果:
\[ \begin{align} \eta_{ktq} = logit^{-1}&(\frac{\bar{\theta}'_{kt}- \beta_q}{\sqrt{\alpha^2_q + (1.7\sigma_{kt})^2}}).\\ \downarrow&\\ \eta_{ktq} = logit^{-1}&(\frac{\bar{\theta}'_{kt}- (\beta_q \color{red}{+ \delta_{kq}})}{\alpha_q}). \end{align} \]
Dynamic Comparative Public Opinion
复杂程度:
Claasseen 2019 < DCPO < DGIRT
发表:
……
劣质源数据
+
数据收集过程的难复刻性
+
难以追踪的人为操作
目标问题
目标问题: 手动输入错误和结果“靠谱性”
解法: 尽可能地自动化
DCPOtools对数据进行预处理(半自动)DCPO进行数据分析(自动)shinystan诊断convergence(自动)诊断什么和如何理解
目标问题: 社会科学研究仍不够透明
解法: Second-order opening, 不止步于分析,还有DGP
A.k.a., 你怎么知道只有一个潜在变量?
Measurement Invariance (Putnick and Bornstein 2016)
EFA: 数据指向
CFA: 理论指向
潜变量矩阵
\[ \Phi = \begin{bmatrix} \phi_{11} & & \\ 0 & \phi_{22} & \\ 0 & \phi_{23} & \phi_{33} \end{bmatrix} \]
测量矩阵
\[ \begin{bmatrix} x_1\\ x_2\\ x_3\\ x_4\\ x_5 \end{bmatrix} = \begin{bmatrix} 1 & 1 & 0\\ \lambda_{21} & 1 & 0\\ \lambda_{31} & 1 & 0\\ \lambda_{41} & 0 & 1\\ \lambda_{51} & 0 & 1 \end{bmatrix} \begin{bmatrix}\xi_1\\ \xi_2\\ \xi_3 \end{bmatrix} + \begin{bmatrix}\epsilon_1\\ \epsilon_2\\ \epsilon_3\\ \epsilon_4\\ \epsilon_5 \end{bmatrix} \]
偏误矩阵
diag Θσ = diag[var(ε1) var(ε2)…var(ε5)]
a.k.a., 最多能连多少线?
Identified: 当Λ、Φ、Θ存在唯一解
Λ: \(X = f(\xi)\)系数 (a.k.a., loading);
Φ: 潜在变量的协方差矩阵;
Θ: 偏误的协方差矩阵。
t rule:
t < q(q + 1)/2
t: 不可见
q:可见
为保证identifable进行的限制
\[\Sigma(\theta) = \Lambda_X\Phi\Lambda_X' + \Theta_\epsilon,\] Σ: 所有可见指标得协方差矩阵
全信息估计 (Maximum likelihood)
有限信息估计
错误:待分析的矩阵不是正定的
错误:不恰当的解
错误:未收敛
Overall: χ2,结果显著1则说明整体模型可能有问题。
Incremental Fit Indices:将模型与基线模型比较
Absolute Fit Indices
对初中生精神状况的调查: 视觉因素(x1~3) + 阅读因素(x4~6)+ (表达因素:x7~9)
lavaan 0.6-19 ended normally after 35 iterations
Estimator ML
Optimization method NLMINB
Number of model parameters 21
Number of observations 301
Model Test User Model:
Test statistic 85.306
Degrees of freedom 24
P-value (Chi-square) 0.000
Model Test Baseline Model:
Test statistic 918.852
Degrees of freedom 36
P-value 0.000
User Model versus Baseline Model:
Comparative Fit Index (CFI) 0.931
Tucker-Lewis Index (TLI) 0.896
Loglikelihood and Information Criteria:
Loglikelihood user model (H0) -3737.745
Loglikelihood unrestricted model (H1) -3695.092
Akaike (AIC) 7517.490
Bayesian (BIC) 7595.339
Sample-size adjusted Bayesian (SABIC) 7528.739
Root Mean Square Error of Approximation:
RMSEA 0.092
90 Percent confidence interval - lower 0.071
90 Percent confidence interval - upper 0.114
P-value H_0: RMSEA <= 0.050 0.001
P-value H_0: RMSEA >= 0.080 0.840
Standardized Root Mean Square Residual:
SRMR 0.065
Parameter Estimates:
Standard errors Standard
Information Expected
Information saturated (h1) model Structured
Latent Variables:
Estimate Std.Err z-value P(>|z|)
visual =~
x1 1.000
x2 0.554 0.100 5.554 0.000
x3 0.729 0.109 6.685 0.000
textual =~
x4 1.000
x5 1.113 0.065 17.014 0.000
x6 0.926 0.055 16.703 0.000
speed =~
x7 1.000
x8 1.180 0.165 7.152 0.000
x9 1.082 0.151 7.155 0.000
Covariances:
Estimate Std.Err z-value P(>|z|)
visual ~~
textual 0.408 0.074 5.552 0.000
speed 0.262 0.056 4.660 0.000
textual ~~
speed 0.173 0.049 3.518 0.000
Variances:
Estimate Std.Err z-value P(>|z|)
.x1 0.549 0.114 4.833 0.000
.x2 1.134 0.102 11.146 0.000
.x3 0.844 0.091 9.317 0.000
.x4 0.371 0.048 7.779 0.000
.x5 0.446 0.058 7.642 0.000
.x6 0.356 0.043 8.277 0.000
.x7 0.799 0.081 9.823 0.000
.x8 0.488 0.074 6.573 0.000
.x9 0.566 0.071 8.003 0.000
visual 0.809 0.145 5.564 0.000
textual 0.979 0.112 8.737 0.000
speed 0.384 0.086 4.451 0.000
Theory matters!
下雪啦天晴啦
下雪别忘穿棉袄
下雪啦天晴啦
天晴别忘戴草帽
带草帽~~~
—《心中的太阳》
今晴,明80%也晴;
今雪,明60%也雪。
| 明天 | |||
|---|---|---|---|
| θ1 | θ2 | ||
| 今天 | θ1 | 0.8 | 0.2 |
| θ2 | 0.6 | 0.4 |
起始点: [0.5 0.5]
\(S_1 = [0.5\; 0.5]\begin{bmatrix}0.8 & 0.2\\ 0.6 & 0.4 \end{bmatrix}=[0.7\; 0.3];\) \(S_2 = [0.7\; 0.3]\begin{bmatrix}0.8 & 0.2\\ 0.6 & 0.4 \end{bmatrix}=[0.74\; 0.26];\) \(S_3 = [0.74\; 0.26]\begin{bmatrix}0.8 & 0.2\\ 0.6 & 0.4 \end{bmatrix}=[0.748\; 0.252];\) \(S_4 = [0.748\; 0.252]\begin{bmatrix}0.8 & 0.2\\ 0.6 & 0.4 \end{bmatrix}=[0.749\; 0.250].\)
再往下算,会发现概率变化越来越小,趋于稳定(congverged)
目前没有统计办法能够证明一个Markov Chain已经converged
增加Convergence可能的方法
\[G = \frac{\bar{\theta_1}- \bar{\theta_2}}{\sqrt{\frac{s_1}{n_1} + \frac{s_2}{n_2}}}.\]
\[\hat{R}= \sqrt{\frac{\hat{var(\theta)}}{W}}\]
Tip
特征:
用于检验Stationarity
Thinning 并不会提高Chain的运算速度、帮助convergence或增强估测质量
何时调Thin
结构
估计:MLE
诊断
除了CFA,SEM还能做什么
敬请注意:SEM结果仍表现相关关系。
处理Measurement errors;处理非线性变量(hint: GSEM);纳入多层效应;处理缺失值
政治民主(1960,1965)与工业化
lavaan 0.6-19 ended normally after 68 iterations
Estimator ML
Optimization method NLMINB
Number of model parameters 31
Number of observations 75
Model Test User Model:
Test statistic 38.125
Degrees of freedom 35
P-value (Chi-square) 0.329
Model Test Baseline Model:
Test statistic 730.654
Degrees of freedom 55
P-value 0.000
User Model versus Baseline Model:
Comparative Fit Index (CFI) 0.995
Tucker-Lewis Index (TLI) 0.993
Loglikelihood and Information Criteria:
Loglikelihood user model (H0) -1547.791
Loglikelihood unrestricted model (H1) -1528.728
Akaike (AIC) 3157.582
Bayesian (BIC) 3229.424
Sample-size adjusted Bayesian (SABIC) 3131.720
Root Mean Square Error of Approximation:
RMSEA 0.035
90 Percent confidence interval - lower 0.000
90 Percent confidence interval - upper 0.092
P-value H_0: RMSEA <= 0.050 0.611
P-value H_0: RMSEA >= 0.080 0.114
Standardized Root Mean Square Residual:
SRMR 0.044
Parameter Estimates:
Standard errors Standard
Information Expected
Information saturated (h1) model Structured
Latent Variables:
Estimate Std.Err z-value P(>|z|)
ind60 =~
x1 1.000
x2 2.180 0.139 15.742 0.000
x3 1.819 0.152 11.967 0.000
dem60 =~
y1 1.000
y2 1.257 0.182 6.889 0.000
y3 1.058 0.151 6.987 0.000
y4 1.265 0.145 8.722 0.000
dem65 =~
y5 1.000
y6 1.186 0.169 7.024 0.000
y7 1.280 0.160 8.002 0.000
y8 1.266 0.158 8.007 0.000
Regressions:
Estimate Std.Err z-value P(>|z|)
dem60 ~
ind60 1.483 0.399 3.715 0.000
dem65 ~
ind60 0.572 0.221 2.586 0.010
dem60 0.837 0.098 8.514 0.000
Covariances:
Estimate Std.Err z-value P(>|z|)
.y1 ~~
.y5 0.624 0.358 1.741 0.082
.y2 ~~
.y4 1.313 0.702 1.871 0.061
.y6 2.153 0.734 2.934 0.003
.y3 ~~
.y7 0.795 0.608 1.308 0.191
.y4 ~~
.y8 0.348 0.442 0.787 0.431
.y6 ~~
.y8 1.356 0.568 2.386 0.017
Variances:
Estimate Std.Err z-value P(>|z|)
.x1 0.082 0.019 4.184 0.000
.x2 0.120 0.070 1.718 0.086
.x3 0.467 0.090 5.177 0.000
.y1 1.891 0.444 4.256 0.000
.y2 7.373 1.374 5.366 0.000
.y3 5.067 0.952 5.324 0.000
.y4 3.148 0.739 4.261 0.000
.y5 2.351 0.480 4.895 0.000
.y6 4.954 0.914 5.419 0.000
.y7 3.431 0.713 4.814 0.000
.y8 3.254 0.695 4.685 0.000
ind60 0.448 0.087 5.173 0.000
.dem60 3.956 0.921 4.295 0.000
.dem65 0.172 0.215 0.803 0.422