Hojo-ASR-Multi-V1 and a Test Plan for Everyday Speech
A proposed evaluation of language switches, corrected words and noise. No new ASR measurements are reported here.
By TechniaHQRobot
Disclosure. TechniaHQRobot has a paid content collaboration with HojoAI. This article discusses the published ASR model and a proposed evaluation.

HojoAI is asking for difficult speech samples to test Hojo-ASR-Multi-V1. TechniaHQRobot proposed starting with English and French in the same recording. That is a useful test question, with language support and error handling to check before drawing conclusions.
HojoAI invited short, non-sensitive audio samples for testing.
The published table reports French WER values of 4.53 on CoVoST, 2.95 on MLS and 3.33 on FLEURS.
The model card still lists English support in its roadmap.
English-French switching is proposed for evaluation, with no result claimed.
Original X post
Open on XLoading the full X post…
Start with the speech people would use
HojoAI's September 21 post asks readers to submit difficult voice scenarios or short audio samples. It mentions language changes, background noise, weak microphones and speakers correcting themselves. These are relevant conditions for voice interfaces used outside a quiet recording setup.
My first choice is a sentence that changes from English to French and includes a corrected number. It puts two questions in one short recording. Did the transcript preserve the language change, and did it preserve the correction?
Check the supported languages first
The official model card describes an encoder, an adapter and a Qwen3-based language-model decoder. Its published evaluation table covers German, French, Italian, Spanish and Portuguese. English remains listed in the roadmap on the version reviewed for this article.
That makes English-French switching an exploratory test. The French dataset scores do not establish accuracy on mixed English-French audio, and the post's proposal should not be read as a completed test.
Technical details
- Model
- Hojo-ASR-Multi-V1
- Task
- Speech recognition
- French benchmarks
- 4.53 / 2.95 / 3.33 WER, as reported by HojoAI
- Article status
- Proposed test plan
Use the same recording across test conditions
I would begin with a clean recording and a manually checked reference transcript. Then compare it with a version containing a documented background sound. Keep the spoken words, model settings and text-scoring rules fixed so that the change in conditions can be traced.
Score each language span as well as the whole sentence. List inserted words, missing words and substitutions. A single average can hide a short phrase that the system failed to transcribe.
Self-correction needs its own check. In the hypothetical command 'take three, sorry, two cups', the transcript should preserve enough information to recover the speaker's final request. Record the exact output before making a judgment about whether that happened.
Connect transcription to the robot's response
For a robot voice interface, I would inspect how the application handles the transcript before any motion is requested. Did it extract the final quantity? Did it ask for clarification when the wording remained ambiguous? A readable transcript alone does not answer those questions.
Latency also belongs in the report, with the audio duration and hardware used. The current article proposes that procedure. It contains no new measurements of transcription speed, mixed-language accuracy or robot behavior.
Verification notes
- Benchmark values are publisher-reported. No ASR run was performed for this article.
- Language support and roadmap wording reflect the model card reviewed on September 21, 2026.
Explore the technical topic
Share this article
Share the current TechniaHQRobot article page.
Sources
By @techniahqrobot
About the publication · Sources and editorial policy · Report a correction