Easily fine-tune (SFT/RL), evaluate, and deploy Qwen, Gemma, or any open agentic LLM/VLM!
-
Updated
Oct 4, 2026 - Python
Easily fine-tune (SFT/RL), evaluate, and deploy Qwen, Gemma, or any open agentic LLM/VLM!
Evaluate your LLM's response with Prometheus and GPT4 💯
👩⚖️ Agent-as-a-Judge: The Magic for Open-Endedness
Dingo: A Comprehensive AI Data, Model and Application Quality Evaluation Tool
Awesome list of TypeSafe AI Jev use cases: 74 demos ranked by likes, 150+ GitHub repos, limits, cost and API examples. CC0
Inference-time scaling for LLMs-as-a-judge.
A native policy enforcement layer for AI coding agents. Built on OPA/Rego.
[ICLR 2025] xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation
xVerify: Efficient Answer Verifier for Reasoning Model Evaluations
Typed judgment tools for MCP agents. TypeSafe's Jev model as verify, screen, find, classify, rerank, decide, compare, extract, review, gate, and score: the model judges, policy decides auto, review, or escalate.
First-of-its-kind AI benchmark for evaluating the protection capabilities of large language model (LLM) guard systems (guardrails and safeguards)
Evaluation package that allows benchmarking of agentic AIs from various sources and frameworks by producing statistical results which can be compared across different use cases and datasets.
CodeUltraFeedback: aligning large language models to coding preferences (TOSEM 2025)
[ICLR 2026 Oral & ICML 2026] Generative Universal Verifier as Multimodal Meta-Reasoner
This is the repo for the survey of Bias and Fairness in IR with LLMs.
Official code release for the EMNLP 2026 paper "Med-Banana: Learning Quality-Controlled Medical Image Editing from Success-and-Failure Trajectories" — Med-Banana-80K dataset + edit–verify–refine system
Solving Inequality Proofs with Large Language Models.
A set of tools to create synthetically-generated data from documents
(NeurIPS 2025) Official implementation for "MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?"
A design-of-experiments platform for evaluating compound AI systems - find which technique drives quality, by how much, and whether the difference is real.
To associate your repository with the llm-as-a-judge topic, visit your repo's landing page and select "manage topics."