Code and experiments for studying temporal confidence calibration in large language models across temporally grounded question answering benchmarks.