
Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages
Google has released Gemini 3.5 Transcribe , a speech-to-text model for real-time voice interfaces and recorded audio.
AINOVAT
Source
60 articles from this source

Google has released Gemini 3.5 Transcribe , a speech-to-text model for real-time voice interfaces and recorded audio.

Cohere has released Parse ( parse-v5.0 ) , a document parsing model aimed at high-volume enterprise ingestion.

Every agent that writes code needs somewhere to run it.

In this tutorial, we use Anthropic’s claude-protein-binder-design dataset, which contains 1,440 AI-designed miniprotein binders tested against 16 targets.

Google Research and UNSW Sydney have released GlucoFM , a self-supervised foundation model for continuous glucose monitoring .

Z.ai has released GLM-5.3-Flash , the first natively multimodal model in the GLM-5 series and the cheapest capable coding model the lab has shipped.
Alibaba’s Qwen team has released Qwen3.8-Flash-Next , an open-weight multimodal Mixture-of-Experts model built for cost per token.

I read every major model release.

IBM has released Granite 4.2 , a family of open reasoning language models in 3B, 8B, and 30B parameter sizes.

Generalist AI has released GEN-1.5 , a robot foundation model that learns a new physical task from a single demonstration.

A team from Google Research and USC has released Mobility-Embedded POIs (ME-POIs) , a framework that folds aggregate human movement into text-based place embeddings.

The ‘neocloud’ label now covers five companies with very different business models.

In this tutorial, we explore a LabPlot -inspired scientific data analysis workflow in Python while preserving the structure and terminology of LabPlot’s aspect tree, analysis kernels, plotting system, and project model.

Harvey has released Harvey Tenet , its first post-trained model, as a research preview as of today.

Superwhisper has released the S1 family of models : S1-Voice, S1-Language, and S1-mini.

PDFs sit at the end of almost every workflow.

Liquid AI has released DSpark draft model checkpoints for three models in its LFM2.5 family: LFM2.5-1.2B-Instruct , LFM2.5-2.6B , and LFM2.5-8B-A1B .

In this tutorial, we design an end-to-end preference-learning workflow using the Anthropic HH-RLHF dataset and Direct Preference Optimization (DPO).

NVIDIA has released TensorRT Model Connect (TRTMC) in public preview, an open-source project that takes a supported Hugging Face or local checkpoint to end-to-end TensorRT inference in two commands.

google/sam is not Segment Anything.

Cartesia has released Sonic-3.6 , the newest version of its real-time text-to-speech model.

Nous Research has shipped Bot Mode for Hermes Agent , its MIT-licensed open source agent.

ByteDance Seed and Tsinghua AIR have released CUDA Agent, an agentic reinforcement learning system that trains a large language model to write GPU kernels that beat a compiler.

MiniMax released MiniMax-Music3 , an open-weights text-to-music model.

In this tutorial, we develop an end-to-end OCR workflow with docTR and explore how modern document understanding pipelines combine text detection, recognition, geometry, layout analysis, structured extraction, and expor…

DeepSeek released DeepSeek Harness v0.1 in developer preview and published the full source code under the MIT license.

In this tutorial, we implement an end-to-end supervised fine-tuning pipeline for the XYZ-Aquila-SFT dataset, Hugging Face Transformers, PyTorch, and PEFT.
Z.ai just released GLM-5.3 .

Cactus Compute has released Needle 2 , an open 45M-parameter model for tool calling, device use, and structured extraction.

In this tutorial, we build an end-to-end workflow for working with the SupraLabs reasoning corpus .
Google has released Gemini 3.7 Flash , the newest model in its Flash tier, three weeks after Gemini 3.6 Flash.
Yesterday, Liquid AI released LFM2.5-VL-3B .

In this tutorial, we build an end-to-end post-training pipeline for a compact instruction-tuned language model using AllenAI’s Open Instruct framework.

NVIDIA introduced open technologies for building always-on AI agents from systems of specialized models.

Object removal models have improved faster than the metrics used to judge them.

Video production is shifting as social clips, ad creative and film pre-visualization move from cloud to local GPUs.

In this tutorial, we build a complete quantitative backtesting workflow with OctoBot and OctoBot-Script while keeping the environment isolated from Colab’s preinstalled dependencies.

webAI has released TwIL-LM , a two-model family of formal-logic reasoners at 1.7B and 3B parameters .
In this tutorial, we implement an end-to-end MiniMax-H3 video generation workflow using ComfyUI as a headless inference backend.

Meta has released Muse Glimmer , a 30-billion-parameter multimodal model distilled from Muse Spark.

ByteDance’s Seed team has introduced SeedRealtime , a native audio-visual full-duplex LLM.
NVIDIA has released NemotronLabs VoiceChat 11B , an open 11B end-to-end speech-to-speech model for real-time, full-duplex conversation.

LLM applications fail in ways traditional software does not.

In this tutorial, we develop an end-to-end sentiment analysis workflow using the Stanford NLP IMDb Large Movie Review Dataset and compare classical machine learning with parameter-efficient transformer fine-tuning.
Long-running agents accumulate state that no transcript captures.

Long-horizon agents accumulate context faster than they resolve tasks.

In this tutorial, we explore the advanced visualization capabilities of the XY Python library by building interactive, scalable, and extensible charts.

Mistral AI has released Shieldstral 1.0 3B , an open-weights, policy-adaptive multimodal safety classifier that treats content moderation as a single yes/no question rather than a fixed taxonomy of harm categories.

Tencent Cloud has open-sourced TencentDB Agent Memory v2.0 , a team-level memory hub for AI agents.

In this tutorial, we build an advanced multimodal retrieval-augmented generation pipeline with NVIDIA NeMo Retriever .

NVIDIA Labs has open-sourced NOOA (NVIDIA Object-Oriented Agents) , a model-agnostic Python framework for building AI agents.

Microsoft has open sourced code-testing-generator , a polyglot agent that writes unit tests and then proves they work.

Liquid AI released LFM2.5-2.6B , an agentic model that runs entirely on-device.

Cloudflare has released Kitesurf , a stateless web browser built specifically for AI agents.

In this tutorial, we explore adaptive experimentation using Meta’s Ax with the modern Client API.

Prime Intellect has open-sourced Prime Agent , a self-improving coding harness designed around two abstractions, the Recursive Language Model (RLM) and Continual Harness.

SkillOpt is a text-space optimizer developed by a team of researchers from Microsoft, Shanghai Jiao Tong University, Tongji University, and Fudan University.

In this tutorial, we build a complete Bayesian marketing mix modeling workflow using Google Meridian .
Meta AI has released Muse Code (in beta), a terminal coding agent in beta, powered by its new Muse Spark 1.2 model .

NVIDIA has released Alpamayo 2 Super , a 34B-parameter vision-language-action (VLA) mode l for autonomous driving, under an open commercial license.