Embodied AI models
Reading time 7 min readPerceptron

Perceptron Mk1.5 brings audio, video tracking and tool use to embodied agents

Perceptron Mk1.5 adds native audio, video object tracking, web tools and sub-agents while cutting end-to-end latency by up to 4.7x in company tests.

By TechniaHQRobot

Perceptron Mk1.5 launch visual from Perceptron

Introduction

Perceptron released Mk1.5 publicly on September 25 as a multimodal model built for embodied agents. The model accepts text, images, video and audio, then returns text, points, boxes, polygons, clips and object tracks.

Perceptron says the same model can sit inside agent control loops for drones, quadrupeds, smart glasses and phones without retraining for each platform. The model sees the platform's available actions as tools, then selects which actions to call and in what order.

Information verified from official sources available as of September 26, 2026.

Perceptron Mk1.5 capabilities and the evidence published at launch

CapabilityLaunch evidenceDeployment check
Video object trackingLeads 3 of 4 tracking benchmarks measured by PerceptronTrack identity through occlusion and camera motion
Egocentric perception50% better hand localization than the strongest Gemini model measuredRepeat on robot or wearable video from the target camera
AudioNative audio input with audio-visual benchmark resultsTest noise, distance and overlapping events
Tool use+36.1 points on MMSearch with toolsMeasure tool-call accuracy and failure recovery
Sub-agentsParallel task execution supportedMeasure total completion time and duplicated work
LatencyUp to 4.7x faster end-to-end than Mk1 in shown workloadsMeasure full sensor-to-command latency on the target system

Benchmark results are reported by Perceptron. They are not independent TechniaHQRobot measurements.

Video tracking is now a native output

Mk1.5 can emit object tracks as timestamped geometries across a video. This differs from running independent detections on every frame and then adding a separate re-identification pipeline downstream.

Perceptron reports that Mk1.5 leads three of the four tracking benchmarks it measured: Molmo2-Track, Ref-DAVIS17 and ReasonVOS. MolmoPoint-8B leads the MeViS valid_u result in the company's comparison.

For robotics, persistent tracks can help maintain identity when an object moves through a scene, when inventory is carried from one place to another, or when a robot needs to keep a target associated with earlier observations.

Egocentric video is aimed at robot and wearable viewpoints

Perceptron trained Mk1.5 on first-person video and reports gains on hand tracking, subtask segmentation and long-form video questions.

The company says Mk1.5 localizes hands in a first-person frame 50 percent better than the strongest Gemini model it measured. That comparison uses Perceptron's own evaluation setup, including task-specific scoring rules described in the launch post.

For embodied systems, hand localization can be useful for learning from demonstrations, tracking human interaction with objects and segmenting long manipulation sequences into smaller actions.

Audio joins text, images and video

Mk1.5 accepts audio natively. Perceptron says the current release is useful for time-aligned transcription and sound-based event detection, while also stating that it does not claim frontier audio-visual performance in this first release.

That limitation is useful context. An embodied agent can use audio as another observation channel, but deployment still needs tests for background noise, microphone distance, overlapping speakers and events that are visible but not audible.

A robot test should compare task completion with audio disabled and enabled under the same scene conditions. That isolates whether the added modality changes the final behavior.

Tool calls include web search, page reading and user-defined functions

Mk1.5 supports arbitrary functions declared in OpenAI-style JSON Schema. In Perceptron's demo loop, those tools include web search, page reading, reverse image search and zoom.

On MMSearch, Perceptron reports a 36.1-point gain when tools are available. The published comparison uses the same tools, prompt and tool-calling loop for each model across 171 questions.

The model can also spawn sub-agents to split work into parallel branches. That design is relevant to physical systems when perception, lookup, planning or verification can run independently before the controller commits to an action.

The latency claim is measured end to end

Perceptron reports roughly two to five times faster request completion than Mk1, with a maximum measured speedup of 4.7 times across the workloads shown in the post.

The latency measurement runs from request submission to the final token. Perceptron reports medians from three runs on one H100 for chat, image question answering and video workloads.

Those conditions matter for robotics. Model latency is only one part of a control loop. Camera capture, preprocessing, network delay, tool execution and actuator control add their own time. A robot deployment should measure sensor-to-command latency rather than reuse API latency as a control-loop figure.

One model is being used across several embodiments

Perceptron says Mk1.5 has already been deployed on drones, robotic dogs, smart glasses and smartphones. In the examples described in the release, each platform exposes its own control surface as callable tools.

This architecture moves platform differences into the tool layer. A drone can expose flight commands while a quadruped exposes locomotion commands, yet the high-level model can use the same reasoning interface.

The strongest test of that claim would use the same task instruction across several bodies, then report task completion, intervention rate and the amount of platform-specific code needed outside the model.

API details and the measurements developers should watch

Mk1.5 is listed with a 32K-token multimodal context window and API model ID perceptron-mk1.5. Perceptron lists pricing at $0.15 per million input tokens and $1.50 per million output tokens.

For embodied-agent work, token price is only one cost. Video frame selection, audio duration, tool calls, control frequency and failed retries can change the cost per completed physical task.

A useful deployment report would publish complete-task success, sensor-to-command latency, tool-call count, human interventions and total compute cost for the same repeated task.

Limitations and missing information

  • The benchmark results and latency figures come from Perceptron's launch evaluation.
  • The 4.7x maximum speedup is workload-dependent and should not be treated as a fixed improvement for every robot control loop.
  • Cross-embodiment deployment does not mean every platform needs zero integration work. Each platform still exposes its own tools, sensors and control interfaces.

Sources and methodology

Share this article

Share the current TechniaHQRobot article page.

Continue reading

Open the latest robotics reporting, Physical AI analysis and hardware notes.

Browse robotics news
Article by @techniahqrobot