1
A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis
首个多数据集基准,专测LLM智能体在微服务故障诊断中的实战能力,AIOps必看。
arXiv:2606.29193v1 Announce Type: cross Abstract: LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to ev…