Functional Estimation in High-Dimensional and Infinite-Dimensional Models
- Series
- Dissertation Defense
- Time
- Thursday, July 16, 2026 - 13:00 for 1.5 hours (actually 80 minutes)
- Location
- Skiles 006
- Speaker
- Minghao Li – Georgia Institute of Technology – minghaoli@gatech.edu
Zoom link: https://gatech.zoom.us/j/97454744253?pwd=bSRS939RDbV6PLbi7Os88aL8yoa1lG.1
Let $(S,\mathcal{A})$ be a measurable space and let $\mathcal{P}$ be a family of probability distributions on it. Given a Banach space $E$, a mapping $\theta:\mathcal{P}\to E$, and a smooth functional $f:E\to\mathbb{R}$, we consider the problem of estimating $f(\theta(P))$ from i.i.d. observations $X_1,\ldots,X_n\sim P$, where $P\in\mathcal{P}$. We write $f\in C^s(E)$ for a functional of Hölder smoothness $s=m+\rho$, where $m\ge 0$ is an integer and $\rho\in(0,1]$. Our aim is to construct estimators of $f(\theta(P))$ and to study their dependence on the sample size $n$, the smoothness $s$, and the dimension or complexity of the parameter $\theta(P)$.
When $\hat\theta_n$ is a $\sqrt{n}$-consistent base estimator of $\theta(P)$, the plug-in estimator $f(\hat\theta_n)$ is asymptotically efficient in classical low-dimensional models, but in high-dimensional and infinite-dimensional settings its bias is often too large to attain the rate $n^{-1/2}$. Existing bias reduction methods, based on iterated bootstrap or on linear aggregation of plug-in estimators, rely on concentration inequalities for $f(\hat\theta_n)$ that are available only for a limited class of models.
In this dissertation we study a class of estimators $T_f(X_1,\ldots,X_n)$ obtained from a Taylor expansion of $f$ about $\hat\theta_n$, of order determined by $s$, together with a sample split. Their analysis uses only bounds on the moments of the linear and higher order terms of $\hat\theta_n-\theta(P)$, rather than concentration inequalities. For functionals of smoothness $s\ge 1$, we derive upper bounds on the $L_p$-errors of $T_f$ whose dependence on $n$, on $s$, and on the dimension or complexity of the parameter matches the minimax lower bounds we obtain. We also give conditions under which these estimators are asymptotically normal and asymptotically efficient.
We develop these results in three settings: models composed of a large number of independent low-dimensional components, high-dimensional exponential families, and functionals of covariance operators in infinite-dimensional subgaussian models, where the complexity of the model is measured by the effective rank of the covariance operator.