Did the model improve, or did the agent get another retry?
We're building a more rigorous Vedika Bench and an enterprise harness with cloud VMs + multiple models.
The goal: make tools, starting state, retries and outcomes easier to inspect.
Both are in development.
ALT AI-assisted Vedika concept illustration titled “What did the agent actually do?” Arrows connect a task, model and tools, a cloud VM, and an execution record. A loop around the VM is labeled “State · retries · outcome.” The footer says “Designed around the complete run,” “In development” and “Concept schematic.” This is a design concept, not a product screenshot or measured result.