MEND - X Β· MULTI-MODEL INFERENCE MATRIX
Three Intelligence Tiers.
Zero Hallucination.
One generic LLM cannot solve factory downtime. PLCs demand sub-100ms edge speed; complex breakdowns require deep deductive reasoning. MEND-X dynamically routes every query to the exact intelligence tier needed.
< 100ms
Minimum Latency
100%
Manual Grounded
3 Tiers
Dynamic Routing
0.0%
Hallucination Target
TIER 00 Β· RESILIENT MULTI-PROVIDEREDGE DEPLOYABLE

Dynamic Failover: Ollama Cloud β Local β Groq
Our most resilient inference setup. Dispatches queries to Ollama Cloud first for high throughput. If network drops or quotas exhaust, it gracefully cascades to your local Ollama daemon, followed by Groq as emergency backup.
Underlying LLM EngineOllama Cloud API / Local Ollama / Groq Fallback
Average LatencyAdaptive (~400ms Cloud / ~2s Local)
Context Window128,000 tokens
Inference Throughput300+ tokens/sec (Cloud) / 45 tokens/sec (Local)
Recommended HostCloud GPU + Local Edge Hybrid
Engineered Superpowers
Zero downtime with automatic 3-tier round-robin fallback
Prioritizes fast cloud GPU inference when API key is present
Offline continuity with local Ollama runtime
Strips chain-of-thought tokens cleanly for valid JSON output
Target Real-World Queries
> Complex multi-machine error code diagnosis under fluctuating connectivityAuto Round-Robin Handled
> Heavy maintenance procedure synthesis with zero downtime requirementAuto Round-Robin Handled
> Offline-first industrial troubleshooting on factory floor gatewaysAuto Round-Robin Handled
> Automated sensor telemetry and fault code correlationAuto Round-Robin Handled
LIVE INFERENCE ROUTER
How MEND-X Decides the Tier
Click a real maintenance scenario below to see the heuristic complexity analyzer evaluate the query and dynamically activate the optimal model.
ROUTER_TRACE // CLASSIFIER_V2
REALTIME ANALYSISINPUT_PROMPT: "PowerFlex 755 Fault 8: Step-by-step deceleration profile tuning and motor test"
TARGET_MACHINE: Allen-Bradley PowerFlex 755
COMPLEXITY_METRIC: 0.58 / 1.00(Procedural Repair)
DECISION_RATIONALE: Multi-step mechanical maintenance procedure requiring sequential action items, tool specs, and parameter verification.
ROUTED_LLM: Groq LPU / openai/gpt-oss-20bLATENCY: 1.0s β 1.8s
TECHNICAL SPECIFICATIONS
Side-by-Side Comparison
| Metric / Capability | Nord (Tier 01) | Forge (Tier 02) | Apex (Tier 03) |
|---|---|---|---|
| Primary Objective | Sub-100ms Error Code Triage | Multi-Step Repair Sequences | Root Cause & Safety Critical |
| Base LLM Engine | groq/compound-mini (Groq LPU) | openai/gpt-oss-20b (Groq LPU) | openai/gpt-oss-120b (Groq LPU) |
| Response Latency | < 100ms | 1.0s β 1.8s | 2.0s β 3.8s |
| Context Window | 8,192 tokens | 128,000 tokens | 128,000 tokens |
| Edge / Offline Capable | Yes (Local IPC / ONNX) | Yes (Plant Server) | Air-Gapped Private VPC |
| Hallucination Mitigation | Strict RAG Masking | Page Citation Grounding | Refusal Circuit + Thresholds |
| Trigger Threshold | complexity < 0.35 | 0.35 β€ complexity < 0.70 | complexity β₯ 0.70 |
Experience the Inference Routing in Real-Time
Test how MEND-X queries live OEM manuals and streams citation-verified repair protocols to line operators in under 8 seconds.
