Abstract
We establish a lower bound on the minimax expected regret of stochastic bandit convex optimization of -Lipschitz functions on the Euclidean ball. This presents the first nontrivial regret lower bound that grows faster than for this problem, establishing that stochastic bandit convex optimization is fundamentally harder than linear bandits. The hard class of convex functions we construct takes the following form in dimension : for an action , each function is the scaled soft maximum of a "tube", (hyperparameterized by ), and a squared distance function, . Here, is an unknown linear transformation, and is an unknown vector which must be learned to minimize the function. Observations are informative about only when the learner's action lies near the tube determined by , satisfying : thus the learner must either find this tube without knowing , or spend observations learning useful directions of . Formally, our regret analysis exploits this tradeoff by bounding the posterior spread of Fisher information matrices obtained under an adaptive sequence of actions. Together, these ingredients give a sample complexity lower bound of to find an -optimal action, which translates to an regret lower bound. We also extend this lower bound to the unconstrained setting where the action space is .