How to evaluate AI agents in 2026: trajectory and tool-call testing, the top eval tools ranked, LLM-as-judge caveats, …