Skip to content

Hidden Cost of ETF Investing: Liu, T. Zhang & Y. Zhang (2026)

Distilled by claude-sonnet-4-6 · extracted Jun 25, 2026, verified Jun 25, 2026

JEL (IAR-assigned): G12, G14, G23, N22 · assigned from the abstract, not the journal

Full structured metadata (methods, scope, relatesTo, topics, datasets): raw Markdown (.md)

paper-summaryasset-pricingequitiesportfolio-sortfama-macbethpanel-regressionopen-accesscc-bypeer-reviewedunreplicateddata:crsp-mutual-fundsdata:wrdsdata:taqdata:morningstardata:thomson-13f

What this is. The paper’s core results (the overnight vs intraday return differential in the US ETF market, its magnitude, and its mechanism), the three tested hypotheses, and the key estimating equations: enough to know what it found and how, without reading all 15 pages. To replicate or extend it, read the full source at the original.

Decomposing ETF close-to-close mid-quote returns into overnight and intraday components (January 2004 to December 2021, 2,916 US ETFs), the paper documents that overnight returns are significantly positive on average (0.78% per month), while intraday returns are not significantly different from zero (-0.15%), producing a persistent overnight-intraday gap of 0.93% per month. This gap is ubiquitous across asset types (equity, fixed income, and other ETFs) and exchanges (NYSE and NASDAQ). Three candidate explanations are tested and the first two are rejected: the differential is not explained by overnight risk being higher than intraday risk (H1 rejected), nor by information asymmetry driving informed traders to exit at the close (H2 rejected). Instead, the evidence supports H3: the gap is driven by excess retail demand near the market open and by arbitrage constraints that slow correction. ETFs with the highest retail demand and highest arbitrage constraints show a monthly differential nearly six times larger than ETFs in the lowest demand and constraint group (1.99% vs 0.31% per month). Using COVID-19 Economic Impact Payments (EIPs) as an exogenous shock to retail demand, the paper shows that EIP months raise the overnight-intraday return difference by 2.37% per month, providing causal evidence for the retail demand channel.

Prior work by Lou, Polk and Skouras (2019) documents a tug-of-war between overnight and intraday returns for individual stocks; this paper establishes the same pattern in ETFs and identifies the underlying mechanism. Lachance (2021) focuses on ETFs’ high overnight returns from a microstructure perspective; this paper complements her work by focusing on the full overnight-intraday differential and by decomposing the sources. Bogousslavsky (2021) documents the cross-section of intraday and overnight stock returns; this paper extends those findings to ETFs and links them to retail demand and arbitrage supply. Berkman et al. (2012) link retail investor attention and bid-ask bounce to inflated open prices; this paper strips out the bid-ask effect via mid-quote returns and shows the return pattern survives. Boehmer et al. (2021) propose an algorithm to identify retail orders in TAQ; the paper uses their method to measure retail order imbalances near the open.

Magnitudes and significance are as reported; \*\*/\*\*\* = 5%/1%. All returns are in percentages. Locators point into the source PDF.

#ResultLocatorMagnitude
R1The all-ETF market portfolio has a significantly positive overnight return and an insignificant intraday return, producing a 0.93%/month overnight-intraday differentialTable 3, Panel A, p. 7Overnight = 0.784%*** (SE 0.165%), Intraday = -0.150% (SE 0.186%), Overnight-Intraday = 0.933%*** (SE 0.207%)
R2ETFs with the highest retail demand have a significantly larger overnight-intraday differential than low-retail-demand ETFs; the composite retail demand spread is 1.162%/monthTable 4, Panel B (bottom row), p. 8High-minus-low composite retail demand = 1.162%*** (SE 0.137%); individual proxies: max return 0.926%*** (SE 0.188%), retail ownership 0.625%*** (SE 0.078%), retail flow 0.360%*** (SE 0.084%)
R3ETFs with the tightest arbitrage constraints have a larger overnight-intraday differential; the composite arbitrage constraint spread is 0.875%/monthTable 4, Panel C (bottom row), p. 8High-minus-low composite arbitrage constraint = 0.875%*** (SE 0.126%); proxies: AP Concentration 0.424%*** (SE 0.095%), IVol 1.012%*** (SE 0.191%), Bid-Ask Spread 0.634%*** (SE 0.121%), Amihud Illiquidity 0.373%*** (SE 0.115%)
R4Jointly, high retail demand and high arbitrage constraints produce a differential nearly six times larger than the low/low group (1.99% vs 0.31% per month)Table 6, Panel A, p. 10High retail demand / high arbitrage constraint = 1.988%*** (SE 0.281%); low / low = 0.312%** (SE 0.152%); difference = 1.143%*** (SE 0.157%)
R5Fama-MacBeth regressions confirm that retail investor ownership positively predicts the overnight-intraday differential after controlling for risk measuresTable 5, col. 3 and col. 6, p. 9Retail Ownership coefficient = 0.898%*** (SE 0.102) in col. 3 (univariate with controls); 0.833%*** (SE 0.092) in col. 6 (full specification); N = 200,118
R6COVID-19 Economic Impact Payments (EIPs) increase the overnight-intraday differential by 2.37%/month, confirming causal role of retail demandTable 8, col. 2-3, p. 11EIP months: +2.372%*** (SE 0.076); Retail_demand x EIP interaction = 1.370%*** (SE 0.086); sample: January 2020 to December 2021, N = 45,492
R7Retail order imbalances near the market open (not the close) drive the overnight-intraday differential, measured directly from TAQTable 9, col. 1-2, p. 11Market open retail imbalance coefficient = 0.082%*** (SE 0.006) on NDiff; market close retail imbalance coefficient = 0.001 (SE 0.002, insignificant); sample: January 2010 to December 2021, N = 3,975,301

Overall (paper’s conclusion). The convenience of buying ETFs during intraday trading hours comes at a cost: retail investors bid up the opening price and arbitrageurs cannot fully correct this by the end of the day. This hidden cost is economically large (4.2 basis points per day on average, or 0.93% per month), ubiquitous across ETF types and exchanges, and causal: exogenous increases in retail demand during EIP months raise the differential. Investors can reduce the cost by purchasing near the market close, when ETF prices are more efficient due to the AP creation and redemption mechanism restoring pricing accuracy.

The paper tests three competing hypotheses about why overnight returns exceed intraday returns for ETFs; there is no formal structural model.

Hypothesis 1 (H1) — Risk-return trade-off. If holding assets overnight entails greater risk than intraday trading, overnight returns should compensate for that risk. Prediction: the overnight-intraday return difference is positively correlated with overnight risk and negatively correlated with intraday risk.

Hypothesis 2 (H2) — Information asymmetry. Following Slezak (1994) and Hong and Wang (2000), if informed investors trade near the close to realize overnight information advantages, closing prices are discounted, raising overnight returns. Prediction: the differential is positively correlated with the proportion of informed (institutional) investors, so ETFs with more retail investors (less informed) should have a smaller differential.

Hypothesis 3 (H3) — Retail demand and arbitrage constraints. Retail investors, who prefer to trade near the market open (Lou, Polk and Skouras (2019)), create excess demand that temporarily inflates opening prices. Arbitrageurs facing inventory limits, execution costs, and concentration constraints cannot immediately correct this mispricing; it unwinds gradually through the trading day, reducing intraday returns. Prediction: the differential is positively correlated with both retail demand and arbitrage constraints.

The identification strategy for H3 exploits the three rounds of COVID-19 Economic Impact Payments (EIPs) distributed in April/May 2020, December 2020/January 2021, and March/April 2021. EIPs are exogenous government transfer payments that increase household cash and retail participation in ETF markets (Divakaruni and Zimmerman (2024)), creating plausibly exogenous variation in retail demand.

The core methodological contribution is the decomposition of ETF close-to-close returns into overnight and intraday components using average NBBO mid-quotes from the first and last 5-minute intervals of the trading day to minimize microstructure noise (bid-ask bounce, as in Berkman et al. (2012)).

Return construction (pp. 4-5). For ETF $i$ on day $t$, letting Pclose,tiP^{i}_{\text{close},t} be the average NBBO mid-quote during the last five minutes and Popen,tiP^{i}_{\text{open},t} the average mid-quote during the first five minutes, the daily intraday return (equation 1) is:

rintraday,ti=Pclose,tiPopen,ti1(1)r^{i}_{\text{intraday},t} = \frac{P^{i}_{\text{close},t}}{P^{i}_{\text{open},t}} - 1 \tag{1}

The daily close-to-close return adjusting for dividends Divti\text{Div}^i_t and cumulative factors CFACPRti\text{CFACPR}^i_t (equation 2) is:

rclose-to-close,ti=Pclose,ti/CFACPRti+Divti/CFACPRtiPclose,t1i/CFACPRt1i1(2)r^{i}_{\text{close-to-close},t} = \frac{P^{i}_{\text{close},t}/\text{CFACPR}^{i}_t + \text{Div}^{i}_t/\text{CFACPR}^{i}_t}{P^{i}_{\text{close},t-1}/\text{CFACPR}^{i}_{t-1}} - 1 \tag{2}

The daily overnight return (equation 3) is:

rovernight,ti=1+rclose-to-close,ti1+rintraday,ti1(3)r^{i}_{\text{overnight},t} = \frac{1 + r^{i}_{\text{close-to-close},t}}{1 + r^{i}_{\text{intraday},t}} - 1 \tag{3}

Monthly returns are standardized to 21 trading days to make observations comparable across months with different trading-day counts (equations 4 and 5):

rintraday,mi=[tm(1+rintraday,ti)]21/n1(4)r^{i}_{\text{intraday},m} = \left[\prod_{t \in m}(1 + r^{i}_{\text{intraday},t})\right]^{21/n} - 1 \tag{4} rovernight,mi=[tm(1+rovernight,ti)]21/n1(5)r^{i}_{\text{overnight},m} = \left[\prod_{t \in m}(1 + r^{i}_{\text{overnight},t})\right]^{21/n} - 1 \tag{5}

where nn is the number of trading days in month mm. Equal-weighted portfolio overnight-intraday return difference (equation 9, denoted NDdiff\text{NDdiff}) is:

rNDdiff,mp=rovernight,mprintraday,mp=ipwm1i(rovernight,mirintraday,mi)(9)r^{p}_{\text{NDdiff},m} = r^{p}_{\text{overnight},m} - r^{p}_{\text{intraday},m} = \sum_{i \in p} w^{i}_{m-1} \left(r^{i}_{\text{overnight},m} - r^{i}_{\text{intraday},m}\right) \tag{9}

Retail demand proxy. Three proxies are constructed: (i) the maximum daily close-to-close return over the past month (MAX), capturing lottery-seeking behavior; (ii) retail investor ownership (the proportion of ETF shares held by retail investors, estimated as total shares minus institutional 13F holdings, then standardized); and (iii) net fund flow from retail investors (quarterly change in retail holdings divided by total shares). A composite retail demand index averages ranks across available proxies.

Arbitrage constraint proxy. Four proxies are combined: idiosyncratic volatility (standard deviation of CAPM residuals over 12 months), bid-ask spread (average closing bid-ask spread), Amihud illiquidity ratio, and AP concentration (inverse of number of authorized participants). A composite arbitrage constraint index averages ranks across available proxies.

Retail order imbalance (equation 10, p. 11). Using the Boehmer et al. (2021) sub-penny price improvement algorithm to identify retail orders in TAQ:

retail order imbalanceit=buy volumeitsell volumeitbuy volumeit+sell volumeit(10)\text{retail order imbalance}_{it} = \frac{\text{buy volume}_{it} - \text{sell volume}_{it}}{\text{buy volume}_{it} + \text{sell volume}_{it}} \tag{10}

This is computed separately for the first 5 minutes after the open and the last 5 minutes before the close.

Portfolio sorts (R2-R4). At the beginning of each month tt, ETFs are sorted into three groups (bottom 30%, middle 40%, top 30%) based on each lagged proxy. Equal-weighted portfolios are held throughout month tt. The reported monthly overnight-intraday return differentials are averaged over the 2004-2021 sample. Standard errors follow Newey and West (1987) with three lags.

Fama-MacBeth cross-sectional regressions (R5, Table 5). Each month, the cross-sectional regression:

NDdiffi,t=αt+βXi,t1+εi,t(FM)\text{NDdiff}_{i,t} = \alpha_t + \boldsymbol{\beta}' \mathbf{X}_{i,t-1} + \varepsilon_{i,t} \tag{FM}

is run with Xi,t1\mathbf{X}_{i,t-1} including Day Risk, Night Risk (or Day Beta, Night Beta), Retail Ownership, Beta, Log Cap, Turnover, and Momentum. Time-series averages of β^t\hat{\beta}_t are reported with Newey-West standard errors (3 lags). Sample: 200,118 monthly observations, January 2004 to December 2021.

Panel regression with benchmark-time fixed effects (R5 robustness, Table 7). To control for underlying asset fundamentals and test H3 jointly:

NDdiffi,t=α+β1Retail_demandi,t1+β2Arbi_constrainti,t1+γZi,t1+μj×t+εi,t(PR)\text{NDdiff}_{i,t} = \alpha + \beta_1 \text{Retail\_demand}_{i,t-1} + \beta_2 \text{Arbi\_constraint}_{i,t-1} + \boldsymbol{\gamma}' \mathbf{Z}_{i,t-1} + \mu_{j \times t} + \varepsilon_{i,t} \tag{PR}

where μj×t\mu_{j \times t} is a benchmark jj by time tt fixed effect (one cell per benchmark-month pair), Zi,t1\mathbf{Z}_{i,t-1} includes Beta, Log Cap, Turnover, and Momentum. Standard errors are clustered at the benchmark level. Sample: 42,700 monthly observations, ETFs with at least two peers sharing the same benchmark.

EIP causal identification regression (R6, Table 8). Restricting to January 2020 to December 2021 and comparing EIP months to non-EIP months in the same pandemic window:

NDdiffi,t=α+β1Retail_demandi,t1+β2EIPt+β3(Retail_demandi,t1×EIPt)+β4Arbi_constrainti,t1+γZi,t1+εi,t(EIP)\text{NDdiff}_{i,t} = \alpha + \beta_1 \text{Retail\_demand}_{i,t-1} + \beta_2 \text{EIP}_t + \beta_3 (\text{Retail\_demand}_{i,t-1} \times \text{EIP}_t) + \beta_4 \text{Arbi\_constraint}_{i,t-1} + \boldsymbol{\gamma}' \mathbf{Z}_{i,t-1} + \varepsilon_{i,t} \tag{EIP}

where EIPt=1\text{EIP}_t = 1 for months in which US households receive EIP payments (April and May 2020, December 2020 and January 2021, March and April 2021). Standard errors are clustered at the fund level. N = 45,492.

Retail order imbalance regression (R7, Table 9). Daily regressions of the overnight-intraday return difference on retail order imbalances near the open and close, with fund and benchmark-time fixed effects, sample January 2010 to December 2021 (excluding 2016-2018). The positive and significant coefficient on open-market imbalance and the insignificant coefficient on close-market imbalance confirm that the retail demand channel operates through open-price inflation, not close-price deflation.

DatasetRole in paperWiki page
CRSP Mutual Fund (CRSPMF) databaseETF identifier, fund metadata, NAV, total net assets, quarterly holdings, inception date, investment style codesCRSP Mutual Funds
CRSP Daily Stock (CRSPSTOCK)Daily open, high, low, close prices; trading volume; shares outstanding; return adjustment factorsWRDS
TAQ databaseNBBO mid-quotes (5-minute intervals at open and close); retail investor order imbalances via Boehmer et al. (2021) algorithmTAQ
MorningstarETF benchmark identifiers, authorized participant (AP) lists, benchmark-level performanceMorningstar
Thomson Reuters s34 filingsQuarterly institutional holding shares; used to construct retail investor ownership as total minus institutionalThomson Reuters 13F

Sample: 2,916 unique US ETFs, January 2004 to December 2021 (217 months). Final sample covers approximately 98% of net assets invested in the US ETF market at end-2021. TAQ analysis restricted to January 2010 to December 2021 (excluding 2016-2018); benchmark panel restricted to ETFs followed by at least two ETFs sharing the same benchmark.

Read the original if you are: (i) building on the return decomposition methodology (equations 1-9, pp. 4-5) for ETF or fund research; (ii) studying the role of retail investors in ETF pricing or market microstructure; (iii) using the TAQ-based retail order imbalance measure following Boehmer et al. (2021) in an ETF context; (iv) working on the cost of ETF investing (the paper’s Online Appendix contains daily-return robustness, value-weighted results, and equity-only subsamples at Tables OA1-OA4); or (v) designing a study that exploits COVID-19 EIPs as an instrument for retail demand shocks.

Source: peer-reviewed, Journal of Banking and Finance vol. 185, article 107621 (2026). This distillation was extracted by an LLM on 2026-06-25 and is not human-verified or independently reproduced. The CC BY 4.0 licence permits mirroring; the verbatim PDF is not hosted in this batch.

Attribution (CC BY 4.0). Liu, Xin, Tianyao (Terry) Zhang, and Yaodong Zhang. “A hidden cost of ETF investing: Retail demand shocks and limits to arbitrage.” Journal of Banking and Finance 185 (2026): 107621. DOI: 10.1016/j.jbankfin.2025.107621. (c) 2026 The Authors. Published by Elsevier B.V. Licensed under CC BY 4.0. This page is an adaptation by the Institute for Automated Research: core results extracted and re-expressed; changes were made.

Found an error or want a topic covered? Open an issue, use the Edit page link above, or email contact@instituteforautomatedresearch.org. Edits are reviewed before publishing; provenance and accuracy are the point.