@cloudkinetix/bmad-enhanced
Version:
Cloud-Kinetix enhanced fork of BMAD-METHOD - Breakthrough Method of Agile AI-driven Development with robust versioning and unified validation.
74 lines (49 loc) • 3.72 kB
Markdown
# Research-Driven Evaluation Suite Design
This task guides the design and implementation of comprehensive evaluation suites for LLM agents through research-driven methodology, focusing on discovering current evaluation best practices rather than prescriptive static implementations.
## Research-First Evaluation Assessment
[[LLM: Begin by researching current AI evaluation methodologies, frameworks, and industry standards. Understand the specific evaluation requirements and assessment landscape before implementing evaluation solutions.]]
### 1. Research Evaluation Approaches
**Evaluation Framework Research Areas**:
- Current LLM agent evaluation frameworks and methodologies (HELM, EleutherAI, etc.)
- Latest developments in LLM evaluation and benchmarking
- Industry-standard evaluation metrics and testing techniques
- Best practices for automated testing of conversational AI systems
- Human evaluation methodologies and quality assessment approaches
**Testing Methodology Research**:
- Evaluation suite design patterns for AI applications
- Automated testing frameworks for LLM applications (PromptFoo, LangSmith, etc.)
- Quality metrics and scoring methodologies for LLM agents
- Regression testing approaches for AI system evaluation
- A/B testing strategies for LLM agent performance comparison
### 2. Research-Based Implementation Strategy
[[LLM: Based on your research findings, implement evaluation suites using current best practices. Focus on:
1. **Evaluation Framework Selection**: Choose evaluation frameworks based on researched capabilities and project requirements
2. **Test Design Strategy**: Design evaluation tests using current assessment methodologies and quality metrics
3. **Automation Implementation**: Implement automated testing using research-backed evaluation tools
4. **Metrics Definition**: Define evaluation metrics using current industry standards and best practices
5. **Validation Methodology**: Establish validation approaches using research-informed evaluation techniques
Document your evaluation implementation choices and rationale based on the research conducted.]]
### 3. Evaluation Suite Framework
**Research Current Assessment Approaches**:
- Investigate automated evaluation techniques for LLM agent quality
- Study human evaluation methodologies for conversational AI systems
- Research performance benchmarking approaches for LLM applications
- Analyze quality scoring and ranking methodologies for LLM agents
**Implementation Areas**:
- Establish evaluation baselines using researched assessment methodologies
- Configure automated testing scenarios based on current evaluation patterns
- Set up quality metrics using research-informed scoring approaches
- Implement validation testing using current best practices
### 4. Validation and Continuous Improvement
**Research Validation Methodologies**:
- Investigate evaluation validation techniques for AI systems
- Study continuous evaluation approaches for production AI applications
- Research evaluation optimization methodologies based on assessment data
- Analyze evaluation cost optimization strategies
**Implementation Validation**:
- Apply research-backed evaluation methodologies to validate assessment effectiveness
- Use current analysis techniques to optimize evaluation coverage and accuracy
- Implement evaluation monitoring based on researched best practices
- Establish evaluation improvement processes using current optimization patterns
---
**Note**: This task emphasizes research-driven evaluation suite design over prescriptive static implementations. Always research current AI evaluation standards and adapt to your specific agent architecture and evaluation requirements.