Long memory (long-range dependence) refers to a stochastic process whose autocorrelations decay hyperbolically rather than geometrically, so that their sum diverges and the spectral density is unbounded at frequency zero. The fractional difference parameter d (equivalently, Hurst exponent H=d+21) quantifies the degree of dependence: d=0 is short memory, 0<d<21 is stationary long memory, d=21 is the boundary, and d=1 is the I(1) unit root. The autoregressive fractionally integrated moving average (ARFIMA) class nests all of these as special cases.
Key Ideas
Long memory generalizes the I(0)/I(1) dichotomy: a series can be integrated of any real order d.
In the time domain, autocorrelations decay as j2H−2 (hyperbolically); in the frequency domain, the spectral density diverges as ∣λ∣1−2H near zero.
ARFIMA(p,d,q) provides the standard parametric model; semi-parametric estimators (Geweke-Porter-Hudak (GPH), local Whittle) estimate d robustly without fully specifying the autoregressive moving average (ARMA) part.
Fractional cointegration: two I(d) series may share a common trend such that their linear combination is I(d−b) for some 0<b≤d.
Structural breaks and regime switching can generate hyperbolic autocorrelation decay in finite samples, mimicking genuine long memory (Diebold-Inoue 2001, Granger-Hyung 2004); disentangling the two is an active research challenge.
How It Works
Formal Definitions
Time domain (McLeod-Hipel 1978): a stationary process has long memory if ∑j=−nn∣ρj∣→∞ as n→∞; equivalently ρj/(cρj2H−2)→1 for some 0<cρ<∞ and H∈(21,1).
Frequency domain: spectral density f(λ)∼cf∣λ∣1−2H as λ→0, i.e., an unbounded pole at the origin. Both definitions are connected via the Hurst exponent H; long memory occurs when H∈(21,1).
Fractional White Noise
yt=(1−L)−dεt
where L is the lag operator, d=H−21, and εt is white noise (Granger 1980, Hosking 1981). Properties:
Stationary for d<21; invertible for d>−21.
Autocorrelations proportional to j2d−1 for large j (hyperbolic decay for d>0).
Spectral density proportional to ∣λ∣−2d near zero.
ARFIMA(p,d,q)
ϕ(L)(1−L)d(yt−μ)=θ(L)εt
introduced by Granger and Joyeux (1980) and Hosking (1981), where ϕ(L) and θ(L) have all roots outside the unit circle and εt is white noise.
Covariance stationary for ∣d∣<21.
Mean-reverting for d<1.
For d>21: infinite variance (but finite-sample implementations use autoregressive integrated moving average (ARIMA) initial conditions).
ARFIMA(0,d,0) reduces to fractional white noise; ARFIMA(p,0,q) reduces to ARMA.
Regress the log-periodogram on log-frequency near zero:
logI(ωj)=a+blog(ωj)+uj
for j=1,…,m, where ωj=2πj/T are Fourier frequencies. The ordinary least squares (OLS) estimate −b^/2 gives d^; asymptotically normal with standard error π(24m)−1/2. Bandwidth m=T is the common rule of thumb. A larger m reduces standard error but induces bias (the log-periodogram approximation holds only near zero); a smaller m reduces bias but inflates variance. Robinson (1995a) refines the estimator to improve small-sample properties.
Local Whittle Estimator (Robinson 1995b)
Minimizes a frequency-domain approximate Gaussian log-likelihood averaged over a neighborhood of zero:
Q(d)=log(m−1j=1∑mλj2dI(λj))−(2d)m−1j=1∑mlogλj
Asymptotically normal and more efficient than GPH. Extended to non-stationary d>21 by Velasco-Robinson (2000).
R/S and Modified R/S Statistics
R/S (Hurst 1951): range of partial-sum deviations from mean, divided by sample standard deviation:
Qn=SnRn=Snmax1≤k≤n∑j=1k(Xj−Xˉn)−min1≤k≤n∑j=1k(Xj−Xˉn)
Under long memory, Qn/nH converges to a constant >21. Under short memory, H=21.
Modified R/S (MR/S) (Lo 1991): replaces Sn with a Newey-West long-run standard deviation estimate σ(q)=[c0+2∑j=1qwj(q)cj]1/2, where wj(q)=1−j/(q+1) are Bartlett weights. Controls for short-memory autocorrelation and heteroscedasticity; the standard R/S statistic is biased toward detecting long memory in the presence of short-memory dependence.
FIGARCH: Long Memory in Conditional Variance (Baillie-Bollerslev-Mikkelsen 1996)
The GARCH(p,q) (generalized autoregressive conditional heteroskedasticity) conditional variance equation can be rewritten as an ARMA in εt2:
{1−α(L)−β(L)}εt2=ω+{1−β(L)}vt
When {1−α(L)−β(L)} has a unit root the process is integrated GARCH (IGARCH; infinite persistence). Fractionally integrated GARCH (FIGARCH) replaces that unit root with a fractional difference:
ϕ(L)(1−L)dεt2=ω+{1−β(L)}vt,0<d<1
yielding impulse-response weights in the conditional variance that decay hyperbolically at rate jd−1 rather than exponentially. Properties:
d=0: reduces to stable GARCH (geometric decay).
d=1: reduces to IGARCH (unit root, infinite persistence).
0<d<1: intermediate — shocks to variance are persistent but eventually die out; conditional variance is covariance nonstationary but finite-variance impulse responses.
Applied to daily USD exchange rates by Baillie-Bollerslev-Mikkelsen (1996): estimated d^≈0.4–0.6, substantially better log-likelihood than GARCH and more parsimonious than IGARCH.
Caveat: The original FIGARCH parameterization can generate negative impulse-response weights for some parameter combinations, which violates non-negativity of the conditional variance. Later modifications (e.g., truncated FIGARCH, HYGARCH) address this constraint.
Fractional Cointegration
If yt∼I(d) and xt∼I(d), they are fractionally cointegrated of order (d,b) if zt=yt−βxt∼I(d−b) for 0<b<d. The error correction representation (Granger 1986):
Φ(L)(1−L)dyt=−γ(1−(1−L)b)zt−1+c(L)εt
Testing for fractional cointegration reduces to testing for fractional integration in the estimated residuals. Dittmann (2004) proposes a three-step error correction model (ECM) estimator for fractionally cointegrated systems.
Sign Tests for d (Delgado-Velasco 2005)
Standard tests for the long-memory parameter (Tanaka 1999 score test, Dickey-Fuller for the unit root boundary d=1) require finite variance innovations and lose all power under stable or heavy-tailed errors. Delgado and Velasco (2005) construct tests based only on the signs of fractionally filtered residuals, which are Rademacher (±1 equiprobable) under any symmetric continuous H0 distribution.
where St(d0)=sign(ε^t(d0)) and ε^t(d0)=(1−L)d0ut.
Under H0, {St(d0)} are i.i.d. ±1 regardless of the innovation distribution, so Tn has an exact distribution computable by Monte Carlo from 2n equi-probable sign sequences. Theorem 1 proves Tn is locally most powerful (LMP) among sign tests: it maximizes the derivative of power with respect to d at d0.
Asymptotic distribution (slow convergence):
n−1/2Tn(d0)→dN(0,6π2)under H0
but convergence is very slow — at n=2000 the empirical size is still only 4.8% vs. nominal 5%. Exact Monte Carlo critical values are always needed in practice.
Composite ARFIMA test (Section 4): replace unknown ARMA parameters with n-consistent robust estimates (least absolute deviations (LAD) or sign-based), then form the projected score statistic Lnpdχ2(1). The test remains valid under infinite-variance innovations, which rule out standard maximum likelihood estimation (MLE).
Monte Carlo comparison (H0: d=1; alternatives d=1−a/n):
Test
N(0,1)
t4
t2
t1 (Cauchy)
Sign (exact)
8.3%
9.1%
10.7%
16.2%
Dickey-Fuller
6.4%
6.3%
5.5%
3.2%
Tanaka
6.0%
5.6%
4.3%
1.6%
(Power at n=400, a=0.5, 5% level. DF and Tanaka use asymptotic critical values; sign test uses exact.)
Sign test power increases with heavier tails while DF and Tanaka lose power — under Cauchy errors DF and Tanaka are essentially useless (below nominal size), while the sign test maintains validity.
Practical guidance: Use the sign test as a robustness check when innovation distributions may be fat-tailed or when infinite variance is plausible (finance, telecommunications, hydrology). Always use exact or Monte Carlo quantiles rather than the normal approximation.
Interface with Structural Breaks
Diebold-Inoue (2001) show that stochastic regime switching (Markov-switching, random-level-shift, and permanent-break models) satisfies the variance-of-partial-sums definition of long memory for any d>0 in finite samples. Granger-Hyung (2004) show a linear process with occasional mean shifts generates sample autocorrelations consistent with ARFIMA(0,d,0). Conversely, Hillebrand (2004) shows ignored GARCH regime changes cause estimated α+β→1, the variance analogue. These results mean any finding of long memory should be preceded by a structural break test; see Structural Break Testing.
Why It Matters
ARFIMA is more parsimonious than multiple unit-root or highly-parameterized AR models for describing gradual decay in autocorrelation.
The d parameter has direct economic content: d close to 21 implies slow mean reversion (long business cycle persistence in inflation, real interest rates); d near 0 implies rapid decay.
Semi-parametric estimators (GPH, local Whittle) are robust to ARMA misspecification, making them practical screening tools.
Fractional cointegration allows for weaker long-run equilibria than classical I(1) cointegration, appropriate when adjustment is extremely slow.
Long memory in variance (FIGARCH, long-memory stochastic volatility) has implications for option pricing and risk management beyond standard GARCH models.
Open Questions
No consensus on optimal bandwidth m for GPH/local Whittle; data-driven selection rules exist but have not converged.
Distinguishing genuine long memory from breaks or regime switching in finite samples remains unresolved; see Structural Break Testing.
Panel and multivariate generalizations are less mature than univariate tests.
Bayesian inference for ARFIMA models is computationally demanding; Doornik-Ooms (2003) exact MLE requires O(n2) computation.