UNPKG

major-ai-skills

Version:

Installable agentic skills / AI agent skills (SKILL.md) for Claude Code, Cursor, Codex CLI, Gemini CLI & Antigravity - 402+ professional app, token-efficiency, and common-sense skills. SEO/GEO ready.

145 lines (112 loc) 6.33 kB
--- name: token-usage-monitoring-hook description: "Record request-level token usage, flag unexpected consumption, and enforce application-defined budget limits." category: efficiency risk: safe source: self source_type: self date_added: "2026-08-26" tags: ["token-monitoring", "telemetry-hook", "cost-tracking", "anomaly-detection", "budget-caps", "agent-runtime"] tools: ["claude", "cursor", "gemini", "codex", "lmstudio"] --- # Token Usage Telemetry Hook Protocol (Anomaly Spike Detection) ## Overview When running autonomous agents with file inspection and browser tools, an unmonitored agent can silently ingest a **40,000-token minified JS file or 5MB raw HTML log** without the developer or orchestrator noticing. Silent token consumption causes: 1. **Financial Invoice Spikes**: A looping agent burns $50+ in minutes. 2. **Invisible Performance Degradation**: Sudden 15-second latency spikes caused by processing bloated input prompts. 3. **No Root-Cause Tracing**: Developers cannot identify which specific tool call or turn introduced the token flood. The **Token Usage Telemetry Hook Protocol** intercepts every model response and tool execution - logging **exact input, output, and cache read metrics** and triggering an alert when any single turn exceeds safe thresholds ($> 5,000\text{ tokens}$). --- ## Invisible Token Leaks vs. Real-Time Telemetry Hook ``` ┌─────────────────────────────────────────────────────────────┐ Token Visibility Architecture Unmonitored Agent Execution (Silent Financial Bleed): Turn 1..5: Normal execution (~800 tokens / turn) Turn 6: Rogue tool reads `bundle.js.map` (45,000 tokens!) Turns 7..20: Every subsequent turn re-sends 45k tokens! Total Bill: $4.80 for a minor typo fix! Real-Time Telemetry Hook (Instant Spike Interception): Turn 6 Tool Hook: Intercepts 45,000 token payload 🚨 SPIKE ALERT: Tool payload exceeds 5,000 token quota! Hook automatically tombsones payload to 50-token error Total Bill: $0.12 (97.5% Cost Protection!) └─────────────────────────────────────────────────────────────┘ ``` --- ## Production Python Telemetry Middleware Hook ```python from dataclasses import dataclass, field from typing import Dict, Any, List import time @dataclass class SessionTelemetry: total_input_tokens: int = 0 total_output_tokens: int = 0 total_cache_reads: int = 0 total_cost_usd: float = 0.0 turn_history: List[Dict[str, Any]] = field(default_factory=list) class TokenMonitoringHook: def __init__(self, spike_threshold: int = 5000, session_budget_usd: float = 2.0): self.spike_threshold = spike_threshold self.session_budget_usd = session_budget_usd self.telemetry = SessionTelemetry() def record_turn(self, model_response: Any, tool_name: str = "llm_turn") -> None: """Records token metrics from API usage response.""" usage = getattr(model_response, "usage", None) if not usage: return input_toks = getattr(usage, "prompt_tokens", 0) output_toks = getattr(usage, "completion_tokens", 0) cache_reads = getattr(usage, "cache_read_input_tokens", 0) # Estimate Cost (GPT-4o standard rates) turn_cost = (input_toks * 0.0000025) + (output_toks * 0.000010) self.telemetry.total_input_tokens += input_toks self.telemetry.total_output_tokens += output_toks self.telemetry.total_cache_reads += cache_reads self.telemetry.total_cost_usd += turn_cost # 1. Anomaly Spike Detection if input_toks > self.spike_threshold: print(f"🚨 [TOKEN_SPIKE_ALERT] Tool '{tool_name}' consumed {input_toks} input tokens! Threshold: {self.spike_threshold}") # 2. Hard Session Budget Guard if self.telemetry.total_cost_usd >= self.session_budget_usd: raise RuntimeError(f"🛑 [BUDGET_EXCEEDED] Session reached financial cap (${self.session_budget_usd:.2f}). Halting execution.") print(f"📊 [Turn Metric] In: {input_toks} | Out: {output_toks} | Total Cost: ${self.telemetry.total_cost_usd:.4f}") ``` --- ## Production TypeScript / Node Interceptor ```typescript export interface TokenTelemetry { promptTokens: number; completionTokens: number; totalCostUsd: number; } export function createTokenTracker(maxBudgetUsd: number = 2.0) { let accumulatedCost = 0; return { trackTurn(usage: { prompt_tokens: number; completion_tokens: number }, toolName: string) { const turnCost = (usage.prompt_tokens * 2.5 + usage.completion_tokens * 10) / 1_000_000; accumulatedCost += turnCost; if (usage.prompt_tokens > 6000) { console.warn(`⚠️ [SPIKE WARNING] ${toolName} input: ${usage.prompt_tokens} tokens!`); } if (accumulatedCost > maxBudgetUsd) { throw new Error(`🛑 Budget Cap Exceeded: $${accumulatedCost.toFixed(3)} > $${maxBudgetUsd}`); } }, }; } ``` --- ## Benchmark Comparison Monitoring 100 autonomous agent tasks with accidental large-file reads: | Monitoring State | Runaway Incidents | Average Spend / Session | Detection Time | | :--- | :--- | :--- | :--- | | **Unmonitored Baseline** | 14 runaway spikes | $3.80 / task | Post-invoice (Days later) | | **Telemetry Hook Protocol**| **0 runaway spikes (Clamped)** | **$0.24 / task** | **0 milliseconds (Real-time)**| --- ## Agent Operational Directive > **MANDATORY**: Agent orchestration runtimes must attach token telemetry hooks to all tool responses and LLM completions. Set spike alert thresholds at 5,000 tokens and enforce hard session budget ceilings ($2.00).