What are the top relevance evaluation toolkits used to measure and improve the quality of search, recommendation, and AI systems, and how do they compare in terms of evaluation metrics, automation, analytics, scalability, ease of use, and support for machine learning workflows?