simulation.create
shallowai.cloudworldmodel/cloud-world-model · Verify this server
Create a temporary anonymous demo cloud simulation from a list of resources and connections (max 2 active simulations per client, up to 10 resources; the returned simulationId is a short-lived unguessable capability that survives MCP transport teardown, but it is cleaned up when the demo lifetime expires or the simulation is deleted). No API key required for this temporary anonymous demo operation. Built-in scenario workflow: call `scenario.list` and pass a returned card's `id` as `scenarioId` to `simulation.create` for server-side graph expansion. For full control, call `scenario.get` and pass its hydrated `resources` and `connections` arrays instead. These are two alternatives — do not send `scenarioId` with `resources` or `connections`. For the catalog EKS Spot Interruption Migration scenario, you may set `scenarioOverrides: { eksSpotInterruption: { startupSeconds } }` with an integer startupSeconds from 0 through 3600 to test a different readiness deadline without copying the graph; this override requires scenarioId and is mutually exclusive with resources and connections. For the Web App Autoscaling scenario, create with `scenarioId: "web-autoscaling"` and `scenarioOverrides: { webAutoscaling: { includeTrafficRecovery: true } }` to activate the optional Traffic Recovery ramp and observe scale-in. `includeTrafficRecovery` is optional and defaults to false, leaving the phase inactive. Opt in before the first simulation.step because replay identity is finalized when stepping starts. This override also requires scenarioId and is mutually exclusive with resources and connections. `scenario.list` returns graph-free cards with bounded active/optional traffic-phase and retry-workload summaries; it is not a source of resource, connection, traffic-pattern, or failure-injection graphs. Scenario traffic and failure presets are not applied automatically. Use it to start any simulation workflow — either with hydrated resources and connections from scenario.get or your own architecture. Do not use it to modify an existing simulation (use simulation.inject_traffic to change load). For the exact owned typical fit, set appWeight:'typical', location.regionKey:'us-east-2' on all four AWS nodes, one ALB with serviceFamily:'alb' and loadBalancerScheme:'internal', two m5.large compute apps with workload:'crud-typical', appRuntime:'node', appWorkerCount:2, appDbPoolSize:250, one db.r5.large MySQL with workloadDatabaseEngine:'mysql', workloadDatabaseVersion:'8.0', maxConnections:500, ALB→each app→DB connections, minInstances=maxInstances=2, autoscaling:false, traffic 20–300 RPS. Only typical-fit-20, typical-fit-100 and typical-fit-200 tuned the typical-v1-20260927c/6aa574d7ff9d3080b88b221bcd59f7d218ae37f0 fit; 300 is an independent holdout and 500 is diagnostic only. Other typical workloads are modeled, not owned. P99 has distinct per-percentile provenance: latencyP99Basis identifies the owned in-VPC internal-ALB fit, a scaled-from-fit estimate (not directly measured), or an uncalibrated generic model. On the exact healthy lean owned graph at 10–1,000 offered target RPS, P99=max(final P95, 7.021919127633514 + 0.00027013891327780484*T) ms; only 10/100/500 RPS were fit, 1,000 RPS was held out. Lean M5 scaling is not a new measurement; typical/heavy and active failures retain uncalibrated P99. predictionEvidence.latencyP99 has measured 0.65–1.35×, scaled 0.50–1.50× (beyond 1,000: 0.25–2×), or uncalibrated 0.50–2× (beyond: 0.25–3×) assumption bounds centered on final P99. These are not confidence intervals or provider measurements. Historical evidence may omit P99. latencyBasis describes the general modeled latency path; use latencyP99Basis specifically for P99. P99 is diagnostic, not scored. For compute, set characteristics.capacityRps for an explicit per-node RPS ceiling at which CPU reaches ~95%; do not use maxThroughput for that compute contract. Kubernetes rejects capacityRps: set maxThroughput for the total cluster RPS ceiling, or nodePools[].maxThroughput for per-node pool capacity. Omitted compute capacityRps uses the selected catalog tier and can intentionally produce a stressed baseline (for example, the AWS m5.large catalog denominator is 2,000 RPS); for a healthy, capacity-bounded compute experiment, declare an explicit per-node capacity such as 500 RPS. That value is an experiment control, not a universal hardware fact. For OCI flexible compute shapes, pass the documented positive integer characteristics.ocpus explicitly; VM.Standard.E4.Flex accepts 1–64 OCPUs and each OCPU maps to 2 vCPUs. OCPU count establishes capacity dimensions only, not provider-specific performance, throughput, or price. An uncounted flexible shape remains an unverified generic estimate. Check GET /api/prediction/generic-shapes for the catalog-derived generic fallback inventory. New prediction-only GCP standard capacity entries include e2-standard-2/4/8/16/32, n1-standard-1/2/4/8/16/32/64/96, and n2-standard-16/32/48/64/80/96/128; provider specifications establish vCPU/memory dimensions, not CWM performance or pricing. Version 1 predictionEvidence explicitly reports legacyGeneric at the top level and on every appCpuByResource item; its note identifies each generic resource's shape and fallback reason. For generic fixed compute, characteristics.instanceCount accepts integer 1–100 represented VMs; capacity aggregates and CPU is per VM. Do not combine it with autoscaling:true, minInstances, or maxInstances. Aurora Serverless v2 remains limited to 1 with multiAz:false or 2 with multiAz:true. For Aurora Serverless ACU limits, use characteristics.config.minCapacity/maxCapacity or flat characteristics.minCapacity/maxCapacity. The exact AWS database shape with serviceFamily: 'aurora-serverless' and size: 'db.serverless' also accepts flat characteristics.minAcu/maxAcu; those aliases are rejected elsewhere, including at the resource root or inside config. For that shape, multiAz:true with instanceCount:2 creates a separately billable reader (<writer-id>-reader) in another AZ. Inspect returned resources and metrics before using simulation.step to observe modeled failover; no AWS timing guarantee is implied. For database connection budgets, set characteristics.connectionDemand on a database: {mode:'declared',declaredConnections:240} uses that plan-time demand without RPS; {mode:'max',declaredConnections:240,idlePoolFloor:200} takes the maximum of load-derived demand and the declared/floor values; omitted configuration preserves load-derived behavior. declaredConnections and idlePoolFloor are ASSUMPTIONS / plan-time budgets (for example, replicas × per-pod pool size), not observed live DB connections. Set maxConnections to the usable limit you intend to test. Per-database metrics report connectionDemandMode, loadDerivedConnections, declaredConnections/idlePoolFloor, and modeledConnections; cost and DB CPU/latency remain based on existing load-driven behavior. Demand above the usable limit adds a bounded, rule-based pool-saturation error signal; it is not a provider-calibrated rate. Capacity, node-bound, SKU, and autoscaling values supplied through this MCP tool are recorded as agent-supplied in the immutable normalizationReceipt; request responseMode: 'full' to inspect it. Generic GKE telemetry and recovery apply only to worker nodes; the control-plane management fee is cost-only, with no modeled control-plane CPU, API throttling, or cooldown. To bound the autoscaled fleet size, set the top-level maxInstances / minInstances parameters. If you do not set maxInstances, the engine uses the provider default — AWS 50, GCP 15, Azure/OCI/DigitalOcean 10 — which may be much larger than your intended fleet size. The response includes effectiveMaxInstances / effectiveMinInstances so you can confirm the bounds that will be enforced. For a targeted CPU HPA scale-out threshold, send the canonical autoscalingTargetCpu field in this create call (for example, autoscalingTargetCpu: 70 for GKE). The compatible aliases scaleOutCpuThreshold, scaleOutCpuPercent, and autoscaleTargetCpuPercent are also accepted; if more than one is sent, their values must agree. Every create response includes hpaAudit with the supplied field, persisted thresholds, and any provider default. For ECS Fargate CPU-only target tracking, set ecsCpuTargetTracking: true, autoscalingTargetCpu, minInstances/maxInstances, and optional scaleOutCooldownSeconds/scaleInCooldownSeconds with simulationSecondsPerStep (default 1). Inspect autoscalingConfig in the compact response or applicationAutoscalingPolicy in the full response. Latency and throughput do not trigger ECS scaling in this mode. These four TOP-LEVEL fields are simulation-wide — the engine applies one CPU threshold identically to every resource's scale decision by default. To make ONE resource scale at a different CPU target than the rest of the simulation (e.g. a GKE cluster scaling out at 60% while an EC2 fleet in the same simulation scales out at 80%), set characteristics.scaleOutCpuThreshold and/or characteristics.scaleInCpuThreshold on that specific resource instead — the per-resource value wins over the simulation-wide default for that resource only. A misnamed near-miss field nested under characteristics (e.g. targetCPUUtilizationPercentage) is rejected with a 400 explaining the correct field name — it is never silently dropped and defaulted. Responses are compact by default: id, name, status, traffic, and a per-resource summary (id, name, status, cpuPercent, routedRps, availabilityState, isRoutable, and recoveryBlockedReason when provided). Pass responseMode: 'full' to get the complete simulation object instead. During a failure workflow, lower traffic to serviceable levels before calling simulation.recover_resource, then use simulation.step until the recovered resource is healthy. Recovery progress is included when applicable: recoveryProgress.state is parked, cooling_down, or healthy, and its parkWindow/cooldown objects report totalSteps, completedSteps, remainingSteps, target, and requiredSteps. Poll simulation.step until state is healthy, then use simulation.metrics to inspect the resulting state and metrics. No prerequisites. Returns the created simulation's id, which every other simulation.* tool consumes; the new simulation also becomes this session's current simulation, so subsequent per-simulation tools may omit simulationId. The likely next tool is simulation.step to advance time. Do not call api.spec to learn the simulation workflow — the tool descriptions in this session contain everything needed. Authenticate with an API key for unlimited persistent simulations.
1 trials · measured 1 day ago
simulation.create scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against ai.cloudworldmodel/cloud-world-model, measured 6 Oct 2026 under methodology v0.2.0. Every measured component scored 100.
Component breakdown
| Component | Weight | Value |
|---|---|---|
| Reliability | 35% | not applicable |
| Schema integrity | 25% | 100.0 |
| Failure behaviour | 15% | not applicable |
| Latency | 15% | not applicable |
| Concurrency | 10% | not applicable |
Tool details
- Transport
- remote
- Credential class
- self-provisionable
- Input schema
- not declared
- Output schema
- not declared
- Side-effect classification
- unclassified
Score history
| Day | Score | Tier | Methodology |
|---|---|---|---|
| 2026-10-06 | 100.0 | shallow | v0.2.0 |
Probe evidence
| Probe | Outcomes |
|---|---|
| schema_integrity | pass: 1 |
Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.
Embed this score
Available for every tool, scored or not — not a verification perk. Always links back to this page.
[](https://vouch.tools/tools/e83a29fe-17b1-414c-b6ff-4d6a491a4324)