Workspace Project:
Enterprise Cloud Platform
No Active Project Scoped
Please select or create a project workspace to enable benchmarking workflows.
Dashboard
Executive summary of retrieval performance and LLM generation quality
Knowledge Corpus
0 Docs
Evaluation Datasets
0 Sets
Benchmarks Created
0 Runs
Evaluations Run
0 Total
Model and Prompt Performance Matrix
Latency vs Cost Trade-Offs
Retrieval Success Trends (Recall vs MRR)
Distribution of Hallucination Scores
No benchmark execution records found. Index a document corpus, add a dataset, and trigger an evaluation benchmark to generate analytics.
Projects & Workspaces
Configure separate workspaces for isolated prompt benchmarks
Active Workspace Projects
➕ Create Project Workspace
Evaluation Datasets
Upload golden-standard QA spreadsheets to grade LLM responses against
📥 Upload Benchmark Dataset
📊 Active QA Sets
Dataset Details
VALIDQA Pairs Preview
| Question | Ground Truth |
|---|
Documents Ingestion
Chunk and index reference materials into the Qdrant vector corpus
📥 Ingest Reference Manual
🗂️ Indexed Documents
Experiment Runs
MLflow-style auditing panel of evaluations and benchmark parameters
🏆 Evaluation Runs Summary
🔬 Detailed Test Case Auditor
⚖️ Evaluation Metrics & Context
| Metric | Score | Reasoning / Explanation |
|---|
Prompt Testing
A/B test prompt templates, grade faithfulness, and choose the optimal layout
📝 Prompt Configuration A
📝 Prompt Configuration B
⚙️ Benchmark Execution Settings
Select a previously saved prompt to use for both A and B. Leave empty to use inline configurations below.
⏳ Evaluations Running
Comparison Leaderboard
| Prompt Config | Quality Score | Faithfulness | Hallucination | Relevancy | BLEU | Avg Latency (ms) |
|---|
Winner Declared
Model Comparison
Run benchmarks across different models to map cost, latency, and quality distributions
⚙️ Multi-Model Benchmark Configuration
⏳ Model Benchmarks Running
Model Leaderboard
| Model | Provider | Faithfulness | Hallucination | Relevancy | BLEU | Avg Latency (ms) | Total Cost ($) |
|---|
PDF Report Generator
Compile metrics matrices, prompt comparisons, and system recommendations into clean audit PDFs
📝 Compile Benchmark Summary
🗄️ Compiled Reports Vault
System Settings
Configure model endpoints, API keys, and manage custom connections
🔑 API Keys
🤖 Custom Endpoints
🖥️ System Health
🗝️ Provider API Credentials
API credentials are saved locally in your active browser session state to authenticate live GPT/Claude/Gemini benchmark queries.
🤖 Register Custom LLM / RAG Endpoint
Registered LLM Configurations
| Name | Provider | Codename | API Base |
|---|