RAG自動評価ツール
BELL DATA, Inc. · Intelligence & Research
Certification per Microsoft Marketplace.
Evidence tier Source Confirmed · 7 captures on record
What the publisher says
As described on Microsoft Marketplace.
LLM-as-a-Judge による RAG自動評価ツール
「LLM-as-a-Judge による RAG自動評価ツール」は、Azure OpenAI と Azure AI Search を活用した RAG(Retrieval-Augmented Generation)型チャットボットの品質を、検索・回答・応答速度の観点から定量的に評価するためのツールです。
Show the rest of the publisher’s description (47 more lines)
生成AIチャットボットの運用において、「検索は本当に適切か」「回答品質は改善しているのか」「設定変更の効果をどう説明するか」といった課題に対し、
LLMを評価者(Judge)として活用することで、人手に頼らない客観的な評価を実現します。RAG改善を“勘や経験”ではなく、“実験と比較”に基づいて判断できる評価基盤を提供します。
サービスの特長
- LLMを評価者としたRAG自動評価(LLM-as-a-Judge)
本ツールでは、回答生成に利用するLLMとは別に、評価専用のLLMを「第三者評価者」として使用します。
これにより、
*
回答内容の正確性・関連性
*
検索結果(Retrieval)の妥当性
*
応答速度(レイテンシ)
を同一基準で自動評価し、RAG構成の違いによる品質差を客観的に把握できます。
- RAG検索方式・モデル差分の定量比較
Azure AI Search を用いた以下の検索方式に対応しています。
- Vector検索(Similarity Search)
- キーワード検索(BM25)
- Hybrid検索 + Semantic Ranker
- 評価用QAテストケースの自動生成
評価に必要なQAテストケースは、アップロードした文書(PDF / TXT)から自動生成されます。
- 自動生成された評価用QAセット
- 既存FAQや想定問答をJSONで持ち込む評価
の両方に対応し、評価データ作成の手間を大幅に削減します。
モデル、Embedding、検索方式、パラメータの違いによる影響を数値と可視化で比較でき、「どの構成が本当に良いのか」を検証可能です。
ご提供機能
*
文書(PDF / TXT)のアップロードと解析
*
評価用QAセットの自動生成
*
RAGによる回答生成
*
LLM-as-a-Judgeによる自動評価
*
回答品質・検索妥当性・レイテンシの集計
*
評価結果のテーブルおよび可視化
導入支援内容
- ツールの動作環境構築支援(Azureを使用)
活用シーン
*
- Azure AI Search のインデックス設計・検索方式検証
- モデル・Embedding選定の評価
- リリース前の品質チェック
注意事項:本サービスは日本国内でのみご利用いただけます。 This service is only available in Japanese language.
本ツールは「生成AI連携サービス for i」サービスと連動した追加支援ツールです。
評価ルールは弊社基準に基づき実装されたものとなり、評価品質を保証するものではありません。
Preview
4 imagesAgent build and provenance
See the full provenance
The layer-by-layer build, the evidence behind each claim, the risk basis and the cross-marketplace links are open to any account. Some rows are disclosed, some the source leaves Unknown; a free account shows you which.
Compliance
- FedRAMPConfirmedNot listed90%, registry-checkedNo FedRAMP Marketplace entry matched this vendor's domain, checked 2026-08-27registry recordas observed 2026-08-27
Confirmed means matched to a public authoritative registry. Claimed means the vendor or its listing states it, not yet cross-checked. A framework not shown was not found in any source we hold, which is not evidence against it. Not listed means a scoped registry check found no match for this vendor's domain: a No is a scoped registry check, not a compliance judgment. Confidence bands: 95% domain-verified, 90% registry-checked, 80% self-attested, 70% weak signal. Self-attested items marked “vendor's site” are gathered from the vendor's own website and are not verified by us.
Sources
Publisher resources
1 linkLinked repositories
Unknown means this listing does not publish a repository. It is not a statement that the code is closed, and a linked repository is not a claim that the publisher wrote it: the registry computes that relationship privately and does not publish it.
Evidence risk is the share of the build you cannot see before you deploy, not a security rating. Sign in to see the layer-by-layer basis for this band.

