Search for a command to run...
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives