A user reported that their Pro Max 5x quota was exhausted in 1.5 hours despite moderate usage, which was unexpected given their previous experience with the service. The user investigated and found that cache_read tokens may be counting at full rate against the rate limit, negating the cost benefit of prompt caching. This issue has significant implications for users who rely on the service for development work. The user provided detailed analysis and suggested improvements to address the problem.