Tutorials
End-to-end walkthroughs that take you from zero to a working evaluation — running your first evaluation, bringing your own data and scoring logic, and connecting agents from popular platforms.
You can also find ready-to-use, runnable examples in the aigo-integrations GitHub repository.
My First Evaluation
Run a harmful-content evaluation on an OpenAI model end to end.
Run Evaluation From AI Atlas
Download and run a ready-made evaluation package with
lf init --atlas.Convert Your Logs into Datasets
Turn external agent conversation logs into Trace datasets.
Write a Sample-Dependent Scorer
Apply per-sample scoring rules for agent benchmarks like tau-bench and SWE-bench.
LangSmith
Connect a stateful LangGraph agent on the LangGraph Platform.
Dify
Connect a Dify agent-chat app with server-side conversation state.
Claude Managed Agents
Connect a Claude Managed Agents deployment as a custom-inference model.
Azure Foundry
Connect an Azure AI Foundry agent while preserving conversation state.
AWS Bedrock
Connect an AWS Bedrock AgentCore Harness as a custom-inference model.