Long Memory and Fractional Integration

long-memoryarfimafractional-integrationgarchfigarchgphhurstr/spersistencespectral-estimation

Definition

Long memory (long-range dependence) refers to a stochastic process whose autocorrelations decay hyperbolically rather than geometrically, so that their sum diverges and the spectral density is unbounded at frequency zero. The fractional difference parameter dd (equivalently, Hurst exponent H=d+12H = d + \tfrac{1}{2}) quantifies the degree of dependence: d=0d = 0 is short memory, 0<d<120 < d < \tfrac{1}{2} is stationary long memory, d=12d = \tfrac{1}{2} is the boundary, and d=1d = 1 is the I(1)I(1) unit root. The autoregressive fractionally integrated moving average (ARFIMA) class nests all of these as special cases.

Key Ideas

How It Works

Formal Definitions

Time domain (McLeod-Hipel 1978): a stationary process has long memory if j=nnρj\sum_{j=-n}^{n} |\rho_j| \to \infty as nn \to \infty; equivalently ρj/(cρj2H2)1\rho_j / (c_\rho j^{2H-2}) \to 1 for some 0<cρ<0 < c_\rho < \infty and H(12,1)H \in (\tfrac{1}{2}, 1).

Frequency domain: spectral density f(λ)cfλ12Hf(\lambda) \sim c_f |\lambda|^{1-2H} as λ0\lambda \to 0, i.e., an unbounded pole at the origin. Both definitions are connected via the Hurst exponent HH; long memory occurs when H(12,1)H \in (\tfrac{1}{2}, 1).

Fractional White Noise

yt=(1L)dεty_t = (1-L)^{-d} \varepsilon_t where LL is the lag operator, d=H12d = H - \tfrac{1}{2}, and εt\varepsilon_t is white noise (Granger 1980, Hosking 1981). Properties:

ARFIMA(p,d,q)

ϕ(L)(1L)d(ytμ)=θ(L)εt\phi(L)(1-L)^d (y_t - \mu) = \theta(L) \varepsilon_t introduced by Granger and Joyeux (1980) and Hosking (1981), where ϕ(L)\phi(L) and θ(L)\theta(L) have all roots outside the unit circle and εt\varepsilon_t is white noise.

GPH Log-Periodogram Estimator (Geweke-Porter-Hudak 1983)

Regress the log-periodogram on log-frequency near zero: logI(ωj)=a+blog(ωj)+uj\log I(\omega_j) = a + b \log(\omega_j) + u_j for j=1,,mj = 1,\ldots,m, where ωj=2πj/T\omega_j = 2\pi j/T are Fourier frequencies. The ordinary least squares (OLS) estimate b^/2-\hat{b}/2 gives d^\hat{d}; asymptotically normal with standard error π(24m)1/2\pi(24m)^{-1/2}. Bandwidth m=Tm = \sqrt{T} is the common rule of thumb. A larger mm reduces standard error but induces bias (the log-periodogram approximation holds only near zero); a smaller mm reduces bias but inflates variance. Robinson (1995a) refines the estimator to improve small-sample properties.

Local Whittle Estimator (Robinson 1995b)

Minimizes a frequency-domain approximate Gaussian log-likelihood averaged over a neighborhood of zero: Q(d)=log ⁣(m1j=1mλj2dI(λj))(2d)m1j=1mlogλjQ(d) = \log\!\left(m^{-1} \sum_{j=1}^{m} \lambda_j^{2d} I(\lambda_j)\right) - (2d)\, m^{-1} \sum_{j=1}^{m} \log \lambda_j Asymptotically normal and more efficient than GPH. Extended to non-stationary d>12d > \tfrac{1}{2} by Velasco-Robinson (2000).

R/S and Modified R/S Statistics

R/S (Hurst 1951): range of partial-sum deviations from mean, divided by sample standard deviation: Qn=RnSn=max1knj=1k(XjXˉn)min1knj=1k(XjXˉn)SnQ_n = \frac{R_n}{S_n} = \frac{\max_{1\leq k\leq n} \sum_{j=1}^{k}(X_j - \bar{X}_n) - \min_{1\leq k\leq n} \sum_{j=1}^{k}(X_j - \bar{X}_n)}{S_n} Under long memory, Qn/nHQ_n / n^H converges to a constant >12> \tfrac{1}{2}. Under short memory, H=12H = \tfrac{1}{2}. Modified R/S (MR/S) (Lo 1991): replaces SnS_n with a Newey-West long-run standard deviation estimate σ(q)=[c0+2j=1qwj(q)cj]1/2\sigma(q) = \bigl[c_0 + 2\sum_{j=1}^{q} w_j(q) c_j\bigr]^{1/2}, where wj(q)=1j/(q+1)w_j(q) = 1 - j/(q+1) are Bartlett weights. Controls for short-memory autocorrelation and heteroscedasticity; the standard R/S statistic is biased toward detecting long memory in the presence of short-memory dependence.

FIGARCH: Long Memory in Conditional Variance (Baillie-Bollerslev-Mikkelsen 1996)

The GARCH(p,q) (generalized autoregressive conditional heteroskedasticity) conditional variance equation can be rewritten as an ARMA in εt2\varepsilon_t^2: {1α(L)β(L)}εt2=ω+{1β(L)}vt\{1 - \alpha(L) - \beta(L)\}\varepsilon_t^2 = \omega + \{1-\beta(L)\}v_t When {1α(L)β(L)}\{1-\alpha(L)-\beta(L)\} has a unit root the process is integrated GARCH (IGARCH; infinite persistence). Fractionally integrated GARCH (FIGARCH) replaces that unit root with a fractional difference: ϕ(L)(1L)dεt2=ω+{1β(L)}vt,0<d<1\phi(L)(1-L)^d \varepsilon_t^2 = \omega + \{1-\beta(L)\}v_t, \quad 0 < d < 1 yielding impulse-response weights in the conditional variance that decay hyperbolically at rate jd1j^{d-1} rather than exponentially. Properties:

Fractional Cointegration

If ytI(d)y_t \sim I(d) and xtI(d)x_t \sim I(d), they are fractionally cointegrated of order (d,b)(d,b) if zt=ytβxtI(db)z_t = y_t - \beta x_t \sim I(d-b) for 0<b<d0 < b < d. The error correction representation (Granger 1986): Φ(L)(1L)dyt=γ(1(1L)b)zt1+c(L)εt\Phi(L)(1-L)^d y_t = -\gamma\bigl(1-(1-L)^b\bigr) z_{t-1} + c(L) \varepsilon_t Testing for fractional cointegration reduces to testing for fractional integration in the estimated residuals. Dittmann (2004) proposes a three-step error correction model (ECM) estimator for fractionally cointegrated systems.

Sign Tests for d (Delgado-Velasco 2005)

Standard tests for the long-memory parameter (Tanaka 1999 score test, Dickey-Fuller for the unit root boundary d=1d=1) require finite variance innovations and lose all power under stable or heavy-tailed errors. Delgado and Velasco (2005) construct tests based only on the signs of fractionally filtered residuals, which are Rademacher (±1\pm 1 equiprobable) under any symmetric continuous H0H_0 distribution.

Test statistic for H0H_0: d=d0d = d_0

Tn(d0)=nj=1n11jγ~nj(d0),γ~nj(d0)=1nt=j+1nSt(d0)Stj(d0)T_n(d_0) = n\sum_{j=1}^{n-1} \frac{1}{j} \tilde\gamma_{nj}(d_0), \qquad \tilde\gamma_{nj}(d_0) = \frac{1}{n}\sum_{t=j+1}^n S_t(d_0) S_{t-j}(d_0)

where St(d0)=sign(ε^t(d0))S_t(d_0) = \text{sign}(\hat\varepsilon_t(d_0)) and ε^t(d0)=(1L)d0ut\hat\varepsilon_t(d_0) = (1-L)^{d_0} u_t.

Under H0H_0, {St(d0)}\{S_t(d_0)\} are i.i.d. ±1\pm 1 regardless of the innovation distribution, so TnT_n has an exact distribution computable by Monte Carlo from 2n2^n equi-probable sign sequences. Theorem 1 proves TnT_n is locally most powerful (LMP) among sign tests: it maximizes the derivative of power with respect to dd at d0d_0.

Asymptotic distribution (slow convergence):

n1/2Tn(d0)dN ⁣(0,π26)under H0n^{-1/2} T_n(d_0) \overset{d}{\to} \mathcal{N}\!\left(0,\, \frac{\pi^2}{6}\right) \quad \text{under H}_0

but convergence is very slow — at n=2000n=2000 the empirical size is still only 4.8% vs. nominal 5%. Exact Monte Carlo critical values are always needed in practice.

Composite ARFIMA test (Section 4): replace unknown ARMA parameters with n\sqrt{n}-consistent robust estimates (least absolute deviations (LAD) or sign-based), then form the projected score statistic Lnpdχ2(1)L_n^p \xrightarrow{d} \chi^2(1). The test remains valid under infinite-variance innovations, which rule out standard maximum likelihood estimation (MLE).

Monte Carlo comparison (H0H_0: d=1d=1; alternatives d=1a/nd = 1-a/\sqrt{n}):

Test N(0,1)\mathcal{N}(0,1) t4t_4 t2t_2 t1t_1 (Cauchy)
Sign (exact) 8.3% 9.1% 10.7% 16.2%
Dickey-Fuller 6.4% 6.3% 5.5% 3.2%
Tanaka 6.0% 5.6% 4.3% 1.6%

(Power at n=400n=400, a=0.5a=0.5, 5% level. DF and Tanaka use asymptotic critical values; sign test uses exact.)

Sign test power increases with heavier tails while DF and Tanaka lose power — under Cauchy errors DF and Tanaka are essentially useless (below nominal size), while the sign test maintains validity.

Practical guidance: Use the sign test as a robustness check when innovation distributions may be fat-tailed or when infinite variance is plausible (finance, telecommunications, hydrology). Always use exact or Monte Carlo quantiles rather than the normal approximation.

Interface with Structural Breaks

Diebold-Inoue (2001) show that stochastic regime switching (Markov-switching, random-level-shift, and permanent-break models) satisfies the variance-of-partial-sums definition of long memory for any d>0d > 0 in finite samples. Granger-Hyung (2004) show a linear process with occasional mean shifts generates sample autocorrelations consistent with ARFIMA(0,d,0). Conversely, Hillebrand (2004) shows ignored GARCH regime changes cause estimated α+β1\alpha+\beta \to 1, the variance analogue. These results mean any finding of long memory should be preceded by a structural break test; see Structural Break Testing.

Why It Matters

Open Questions

Related