Skip to content

Subjective Performance Evaluation and Influence Activities: de Janvry et al. (2023)

Distilled by claude-sonnet-4-6 · extracted Jun 25, 2026, verified Jun 25, 2026

JEL (IAR-assigned): D73, H83, J45, M54, O17, O18, P25 · assigned from the abstract, not the journal

Full structured metadata (methods, scope, relatesTo, topics, datasets): raw Markdown (.md)

paper-summarybureaucracyincentivespublic-sectorfield-experimentpanel-regressionpeer-reviewedunreplicated

What this is. This page is a distilled skeleton of the paper. Read the original at https://doi.org/10.1257/aer.20211207 to replicate or extend.

De Janvry, He, Sadoulet, Wang, and Zhang (2023) run a randomized field experiment with 3,785 college graduate civil servants (“CGCSs”) in two Chinese provinces. The experiment randomizes whether each civil servant learns the identity of her performance evaluator at the start of the evaluation cycle (the “revealed” scheme, mimicking the status quo) or only learns that one of her two supervisors will be randomly selected as evaluator at the end of the year (the “masked” scheme). Under the revealed scheme the evaluating supervisor gives 0.311 higher assessment score points than the nonevaluating supervisor (0.24 SD; DV SD = 1.31): consistent with the CGCS engaging in evaluator-specific influence activities to improve her evaluation outcome, which determines her promotion to a permanent civil service position. Switching to the masked scheme eliminates this assessment asymmetry and improves multiple performance indicators: colleague assessments rise by 0.22 points on a 7-point scale, nonevaluator assessments rise by 0.22 points, and performance-linked monthly wages increase by roughly 2.3 percent. The results show that a low-cost modification of the evaluation scheme can improve bureaucratic work performance.

#ResultLocatorMagnitude
R1Evaluator gives higher assessment than nonevaluator under revealed schemeTable 2, col 1, p.7810.311 (SE 0.082); 0.24 SD (paper text p.780; DV SD=1.31)
R2Evaluator-nonevaluator asymmetry disappears under masked schemeTable 2, col 2, p.781-0.097 (SE 0.121), not significant
R3Masked scheme increases colleague assessment scoreTable 3, Panel A, col 1, p.782+0.217 on 1-7 scale (SE 0.035)
R4Masked scheme increases probability rated top 10% by colleaguesTable 3, Panel A, col 2, p.782+7.7 pp (SE 1.3)
R5Masked scheme increases nonevaluator supervisor assessmentTable 3, Panel B, col 3, p.782+0.215 on 1-7 scale (SE 0.059)
R6Masked scheme increases performance-linked monthly wageTable 3, Panel C, col 1, p.782+48.81 yuan (~2.3%) (SE 22.41)
R7One-point increase in evaluator score raises promotion probabilityTable 4, col 1, p.787+7.3 pp (SE 1.1)
R8Hometown tie with evaluator raises evaluator assessment under revealed scheme onlyTable 7, Panel A, col 2, p.791+0.189 (SE 0.067); null in masked scheme (-0.067, SE 0.088)

Overall. Evidence from Tables 2-7 consistently supports the existence of evaluator-specific influence activities under the revealed scheme and shows that the masked scheme eliminates such activities while improving actual job performance. The hometown favoritism result (R8) isolates “bottom-up” influence activities from “top-down” evaluator preferences: hometown favoritism appears only under the revealed scheme, where the CGCS knows who the evaluator is and can direct influence efforts accordingly.

The paper has no full structural model; Section II (pp. 777-779) presents a conceptual framework to rationalize the experimental design and derive testable propositions.

A CGCS allocates effort across three types of activity: XX (common productive tasks valued by both supervisors), xjx_j (supervisor-jj-specific productive influence activities, i.e., tasks assigned or observed mainly by supervisor jj), and uju_j (nonproductive influence activities directed at supervisor jj, e.g., personal favors). Following Milgrom and Roberts (1988), xjx_j are “productive influence activities” and uju_j are “nonproductive influence activities.”

The organization’s performance measure uses only productive activities (p.778):

P = X + x_1 + x_2 \tag{1}

Supervisor jj‘s subjective assessment score is (p.778):

Y_j = \alpha X + x_j + u_j, \quad j = 1, 2 \tag{2}

where α>0\alpha > 0 is the relative weight the supervisor places on common productive activities over supervisor-specific influence activities.

Each CGCS maximizes utility subject to a total time constraint of TT (p.778):

\max_{X,\, x,\, u} V = \alpha X + \sum_{j \in \{1,2\}} s_j (x_j + u_j) - G(X) - g\!\left(\sum_j x_j\right) - h\!\left(\sum_j u_j\right) \tag{3}

subject to X+jxj+juj=TX + \sum_j x_j + \sum_j u_j = T, X,xj,uj[0,T]X, x_j, u_j \in [0, T],

where sjs_j is the probability supervisor jj‘s assessment determines the CGCS’s reward (jsj=1\sum_j s_j = 1), and GG, gg, hh are strictly convex cost functions. Under the revealed scheme s1=1s_1 = 1, s2=0s_2 = 0; under the masked scheme s1=s2=1/2s_1 = s_2 = 1/2.

Two propositions follow from solving the CGCS’s maximization problem (pp.778-779):

Proposition 1. Under the revealed scheme, the CGCS engages in evaluator-specific influence activities (xj>0x_j > 0, uj>0u_j > 0), and the evaluating supervisor gives a higher assessment (YjY_j) than the nonevaluating supervisor.

Proposition 2. Compared to the revealed scheme, the masked scheme increases common productive effort (XX) and improves overall work performance (PP). The masked scheme raises the nonevaluator’s assessment unambiguously, but its effect on the evaluator’s assessment is ambiguous (the evaluator benefits from more XX but loses evaluator-specific influence).

The primary identification strategy is random assignment of CGCSs to the revealed vs. masked evaluation scheme. In collaboration with two Chinese provincial governments in 2017, the authors randomized all 3,785 CGCSs employed in that year across 788 townships (Section I.B-C, pp.773-775). Two-thirds were assigned to the revealed scheme and one-third to the masked scheme. Randomization was conducted at the work-unit level; since 83.9 percent of units had only one CGCS, this is statistically nearly equivalent to individual-level randomization.

Each CGCS reports to a party leader and an administrative leader under China’s dual-leadership governance structure (Shirk 1993). One of the two supervisors was randomly selected as evaluator. In the revealed scheme, the CGCS was notified of the evaluator’s identity at the start of the evaluation year. In the masked scheme, the CGCS was told only that one supervisor would be randomly selected at year-end. Neither supervisor was informed of the selection. Official government notifications with formal stamps were sent to all CGCSs to establish credibility.

The benchmark performance measure is the average colleague assessment, collected through anonymous surveys of coworkers who have no incentive to inflate or deflate CGCS evaluations (they are not in the CGCS’s evaluation chain and do not compete with her for promotion). Performance is further benchmarked against both supervisor assessments and administrative salary records verified by the provincial governments.

The method builds on the principal-agent framework of Baker, Gibbons, and Murphy (1994), operationalized as an RCT in the spirit of Finan, Olken, and Pande (2015) on public employee incentives. Lazear and Oyer (2012) survey the theoretical literature motivating the empirical test. Prendergast and Topel (1996) develop the theoretical foundations of favoritism in organizations under subjective evaluation. Wu (2017) provides a related natural experiment varying authority allocation in Chinese media, complementing this paper’s approach of directly cross-randomizing the employee’s knowledge of the evaluator’s identity.

Specification 1 (Proposition 1 test). Using the revealed-scheme subsample, the paper estimates (p.780, eq.1):

\text{Sup1\_Edge}_{icst} = \alpha \times \text{Sup1\_Eval}_i + \gamma_c + \lambda_s + \phi_t + \varepsilon_{icst} \tag{4}

where Sup1_Edgeicst\text{Sup1\_Edge}_{icst} is Supervisor 1’s assessment minus Supervisor 2’s assessment for CGCS ii in county cc, CGCS type ss, cohort tt. Sup1_Evali\text{Sup1\_Eval}_i is a dummy for whether Supervisor 1 is the randomly selected evaluator. γc\gamma_c, λs\lambda_s, ϕt\phi_t are county, CGCS-type, and cohort fixed effects. Standard errors are clustered at the work-unit level. Because the evaluator is chosen randomly, α\alpha causally identifies the additional positiveness of the evaluating supervisor’s assessment due to influence activities. The same regression is estimated on the masked-scheme sample (Table 2, cols 1 and 2, p.781) to check that the asymmetry disappears when evaluator identity is withheld.

Specification 2 (Proposition 2 test). Using the full sample, the paper estimates (p.782, eq.2):

Y_{icst} = \alpha \times \text{Mask}_i + \gamma_c + \lambda_s + \phi_t + \varepsilon_{icst} \tag{5}

where YicstY_{icst} is a performance measure (colleague assessment, supervisor assessment, performance pay) and Maski\text{Mask}_i is a dummy for being assigned to the masked scheme. Same fixed effects and clustering. Random scheme assignment means α\alpha identifies the causal effect of masking on performance (Table 3, p.782).

Balance. Table 1 (p.776) shows no statistically significant differences in CGCS characteristics (age, gender, college type, major, party membership, CEE score, risk aversion, local birth) across the two schemes; the joint F-statistic is 0.90 (p = 0.54).

Robustness. The paper controls for LASSO-selected covariates (online Appendix Tables A7, A13), applies Lee (2009) bounds for non-random attrition (online Appendix Tables A9, A15), and uses an interaction approach on the full sample (online Appendix Table A10). Results are stable across these checks.

DatasetRole in paperWiki page
Author-collected CGCS baseline and endline surveys (Sep 2017, Jun 2018)Primary performance measures: colleague assessments, supervisor assessments, self-assessments, job-task allocation, influence activity proxiesno page yet
Chinese provincial government administrative records (2017-2018)Promotion outcomes (permanent civil service placement) and salary data verified against administrative recordsno page yet

Sample: 3,785 CGCSs (“College Graduate Civil Servants” hired through China’s “3+1 Supports” program) in two provinces (Province A coastal, Province B inland), cohorts admitted 2016 and 2017. Endline: 2,854 CGCSs after 24.5 percent attrition, primarily from reassignment between townships (14.9 percent) and voluntary exits to graduate school or civil service exams (7.4 percent). Position types: township government clerks (poverty alleviation and agricultural support), primary school teachers, and township clinic nurses. Randomization at the work-unit level across 788 townships.

Read Section II for formal proofs of the two propositions and model extensions in online Appendices C-E. Read Section III.A (Table 2, p.781) for the evaluator-asymmetry test. Read Section III.B (Table 3, p.782) for the performance-improvement results and Section III.C (Table 4, p.787) for the promotion-weight evidence confirming the stakes are real. Read Section IV for the mechanism analysis: Table 5 (p.789) for productive influence activities (task reallocation toward evaluator-assigned tasks), Table 6 (p.789) for indirect proxies of nonproductive influence activities, and Table 7 (p.791) for hometown favoritism as a test of bottom-up vs. top-down favoritism. Read Section IV.C-D (pp.791-796) for the full battery of robustness checks ruling out evaluator behavioral change and information-quality alternative explanations.

Useful for: researchers studying subjective performance evaluation, influence activities in bureaucracies, and personnel economics of the public sector; practitioners designing evaluation systems in organizations with multiple supervisors or dual-leadership structures.

This paper is published in the American Economic Review 113(3), 2023 under AEA standard copyright. No CC license was found in Crossref metadata (checked 2026-06-25). Extract-only.

de Janvry, Alain, Guojun He, Elisabeth Sadoulet, Shaoda Wang, and Qiong Zhang. “Subjective Performance Evaluation, Influence Activities, and Bureaucratic Work Behavior: Evidence from China.” American Economic Review 113, no. 3 (March 2023): 766-799. https://doi.org/10.1257/aer.20211207

Replication data: de Janvry et al. (2023). Replication Data for: Subjective Performance Evaluation, Influence Activities, and Bureaucratic Work Behavior: Evidence from China. AEA/ICPSR. https://doi.org/10.3886/E182787V1

LLM-distilled by paper-distiller (claude-sonnet-4-6), 2026-06-25. Not human-verified. Not reproduced.

Found an error or want a topic covered? Open an issue, use the Edit page link above, or email contact@instituteforautomatedresearch.org. Edits are reviewed before publishing; provenance and accuracy are the point.