Terminal-Bench-Science: Evaluating AI Agents on Complex Real-World Scientific Workflows in the Terminal