aiwg
Version:
Cognitive architecture for AI-augmented software development with structured memory, ensemble validation, and closed-loop correction. FAIR-aligned artifacts, 84% cost reduction via human-in-the-loop, standards adopted by 100+ organizations.
231 lines (162 loc) • 9.64 kB
Markdown
# v2026.2.0 - "Universal Deploy" Release
**Released**: February 8, 2026
The largest AIWG release to date. Universal deployment ensures all 8 coding platforms receive all 4 artifact types. External Ralph loops enable crash-resilient multi-session task execution. A complete research framework, agent persistence system, and regression testing capability round out a transformative update spanning 95 commits.
## Highlights
| What Changed | Why You Care |
|--------------|--------------|
| **Universal deployment** | All 8 providers now receive all 4 artifact types — 32 combinations |
| **External Ralph loop** | Crash-resilient iterative task execution across sessions (6-8 hours) |
| **Research framework** | 8 agents, 10 commands, 8 templates for research-backed development |
| **Rules as artifact type** | Enforcement rules propagate to every platform automatically |
| **Agent persistence** | Anti-laziness detection, HITL gates, cross-loop learning |
| **Regression testing** | Automated regression detection integrated across SDLC |
| **Unified extension system** | Phase 4 complete: 10 extension types, 40 CLI commands |
| **GitHub Copilot full support** | Rules and skills deploy alongside agents and commands |
| **Test consolidation** | 31.7% test reduction (3,837 → 2,619) with zero coverage loss |
| **Research-first rules** | Agents must research before decisions, parse instructions before acting |
## Universal Deployment Architecture
The headline feature: **every provider gets everything**.
Previously, different platforms had different levels of support — some got agents, some got commands, but no platform got all artifact types. Now all 8 providers receive all 4 artifact types (agents, commands, skills, rules), with per-provider support levels defining how each platform discovers them.
### Support Matrix
| Platform | Agents | Commands | Skills | Rules |
|----------|--------|----------|--------|-------|
| **Claude Code** | native | native | native | native |
| **OpenAI/Codex** | native | native | conventional | native |
| **GitHub Copilot** | native | native | conventional | native |
| **Cursor** | native | native | conventional | native |
| **Factory AI** | native | native | conventional | native |
| **OpenCode** | native | native | conventional | conventional |
| **Warp Terminal** | aggregated | aggregated | conventional | conventional |
| **Windsurf** | aggregated | native | conventional | conventional |
Support levels:
- **native** — Platform auto-discovers files in its standard directory
- **conventional** — AIWG places files in a consistent location the platform can reference
- **aggregated** — Content aggregated into a single context file (e.g., WARP.md, AGENTS.md)
### Deployment Paths
```
aiwg use sdlc # Deploy to Claude Code (default)
aiwg use sdlc --provider codex # Deploy to OpenAI/Codex
aiwg use sdlc --provider copilot # Deploy to GitHub Copilot
aiwg use sdlc --provider cursor # Deploy to Cursor
aiwg use sdlc --provider factory # Deploy to Factory AI
aiwg use sdlc --provider opencode # Deploy to OpenCode
aiwg use sdlc --provider warp # Deploy to Warp Terminal
aiwg use sdlc --provider windsurf # Deploy to Windsurf
```
### Rules as a First-Class Artifact
Rules are no longer just documentation — they're deployable artifacts that propagate to every platform:
- **Discrete-file platforms** (Claude, Codex, Cursor, Factory, OpenCode, Warp) receive individual rule files in their rules directory
- **Content-injection platforms** (Copilot, Windsurf) receive rules injected into their aggregated context files
- 7 core enforcement rules deploy by default: no-attribution, token-security, versioning, citation-policy, anti-laziness, executable-feedback, failure-mitigation
## External Ralph Loop
Crash-resilient iterative task execution for long-running operations that span multiple sessions.
### The Problem
Internal Ralph (`/ralph`) runs within a single Claude Code session. If the session crashes, context is corrupted, or the task takes 6-8 hours, internal Ralph loses state.
### The Solution
External Ralph wraps Claude Code sessions with an external supervisor:
```bash
# Long-running migration (up to 20 iterations, $5 budget)
/ralph-external "Migrate codebase to TypeScript" \
--completion "npx tsc --noEmit exits 0" \
--max-iterations 20 \
--budget 5.0
# Cross-provider support
/ralph-external "Implement feature X" \
--completion "npm test passes" \
--provider codex
# Research-backed options
/ralph-external "Fix all tests" \
--completion "npm test passes" \
--memory complex \
--cross-task
```
### 4-Layer Intelligent Control
1. **Loop Lifecycle** — Initialization, iteration management, termination
2. **Intelligent Control** — Reflexion memory, iteration analytics, early stopping, best output selection
3. **Cross-Task Learning** — Similar task detection, strategy transfer from past loops
4. **Multi-Loop Management** — Concurrent loop execution, monitoring dashboard
### When to Use Which
| Feature | Internal (`/ralph`) | External (`/ralph-external`) |
|---------|---------------------|------------------------------|
| Session duration | Single session | Multi-session (6-8 hours) |
| Crash recovery | Limited | Full supervisor recovery |
| State capture | Basic | Comprehensive snapshots |
| Context corruption | Risk exists | External state preserved |
## Research Framework
A complete research management system built on FAIR principles and GRADE evidence assessment.
### Components
- **8 Research Agents**: Quality Assessor, Citation Verifier, Writing Validator, Prompt Optimizer, Content Diversifier, and more
- **10 Research Commands**: `/verify-citations`, `/grade-report`, `/citation-check`, `/corpus-health`, `/quality-assess`, etc.
- **8 Research Templates**: Frontmatter, quality assessment, evidence review formats
- **W3C PROV Provenance**: Full derivation tracking for all artifacts
- **GRADE Assessment**: Evidence quality scoring per systematic review methodology
## Agent Persistence and Anti-Laziness
Prevents agents from taking destructive shortcuts when tasks get difficult.
### Detection
The Laziness Detector agent monitors for avoidance patterns:
- Test deletion or skipping (`.skip()`, `xit()`)
- Feature removal instead of fixing
- Assertion weakening (`expect(true).toBe(true)`)
- Error suppression (empty catch blocks)
### Recovery Protocol
When avoidance is detected: **PAUSE → DIAGNOSE → ADAPT → RETRY → ESCALATE**
1. Stop execution, preserve state
2. Analyze root cause (cognitive overload? knowledge gap?)
3. Select recovery strategy (simplify task, change approach, reduce scope)
4. Attempt fix (max 3 attempts)
5. Escalate to human with full context if unresolved
### HITL Gates
Human-in-the-loop gates at every phase transition with rich context display, artifact preview, diff view, and configurable approval workflows.
## Research-First and Instruction-Following
Two new HIGH-enforcement rules address common complaints about AI coding assistants:
### `research-before-decision.md`
Agents must research codebase, docs, and external sources before making technical decisions. No more "yolo through problems without searching first."
### `instruction-comprehension.md`
Agents must parse and confirm understanding of all instructions before acting. Re-read instructions on failure to prevent whack-a-mole loops.
## Test Suite Consolidation
Reduced from ~3,837 to ~2,619 tests (31.7% reduction) with zero coverage loss.
### How
The key insight: `test.each`/`it.each` expands each case as a separate test in vitest. Using `for`/`forEach` inside a single `it()` block consolidates N parameterized cases into 1 test while maintaining all assertions.
```typescript
// Before: 5 separate tests
it.each(['a', 'b', 'c', 'd', 'e'])('validates %s', (v) => {
expect(validate(v)).toBe(true);
});
// After: 1 test covering all 5 cases
it('validates all inputs', () => {
for (const v of ['a', 'b', 'c', 'd', 'e']) {
expect(validate(v)).toBe(true);
}
});
```
## CI/CD Improvements
- Markdown lint job removed (framework content files never conform to strict lint rules)
- TypeScript compilation errors resolved for all Platform record types
- Docker CI compatibility improved with `skipIf` guards for tsx-dependent tests
- Ralph-external race condition fixed (async provider registration properly awaited)
- Flaky timing assertions relaxed for environment-independent execution
## Breaking Changes
None. This release is fully backward-compatible with v2026.1.x.
## Upgrade Path
```bash
# Update via npm
npm install -g aiwg@latest
# Re-deploy frameworks to pick up new rules and universal deployment
aiwg use sdlc
aiwg use marketing # If using marketing framework
# For other providers
aiwg use sdlc --provider <your-provider>
```
## Statistics
- **95 commits** since v2026.1.7
- **8 providers** with universal artifact deployment
- **32 provider × artifact combinations** supported
- **2,619 tests** passing (down from 3,837, same coverage)
- **7 core enforcement rules** deployed to all platforms
- **8 research agents** with 10 commands and 8 templates
- **4-layer** intelligent control system for external Ralph
## Full Changelog
See [CHANGELOG.md](../../CHANGELOG.md) for the complete list of changes.
---
## v2026.2.1 Patch (February 8, 2026)
CI/CD fix: npm publishing to public npmjs.org is now automated on tag push. Previously required manual dispatch with OTP. Uses a separate `NPMJS_TOKEN` granular access token that bypasses 2FA for CI.