openvino_genai.ASRPipeline#
- class openvino_genai.ASRPipeline#
Bases:
pybind11_objectAutomatic speech recognition pipeline
- __init__(self: openvino_genai.py_openvino_genai.ASRPipeline, models_path: os.PathLike | str | bytes, device: str, **kwargs) None#
ASRPipeline class constructor. models_path (os.PathLike): Path to the model file. device (str): Device to run the model on (e.g., CPU, GPU).
Methods
__delattr__(name, /)Implement delattr(self, name).
__dir__()Default dir() implementation.
__eq__(value, /)Return self==value.
__format__(format_spec, /)Default object formatter.
__ge__(value, /)Return self>=value.
__getattribute__(name, /)Return getattr(self, name).
Helper for pickle.
__gt__(value, /)Return self>value.
__hash__()Return hash(self).
__init__(self, models_path, device, **kwargs)ASRPipeline class constructor.
This method is called when a class is subclassed.
__le__(value, /)Return self<=value.
__lt__(value, /)Return self<value.
__ne__(value, /)Return self!=value.
__new__(**kwargs)Helper for pickle.
__reduce_ex__(protocol, /)Helper for pickle.
__repr__()Return repr(self).
__setattr__(name, value, /)Implement setattr(self, name, value).
Size of object in memory, in bytes.
__str__()Return str(self).
Abstract classes can override this to customize issubclass().
generate(self, audio_inputs[, ...])High level generate that receives raw speech as a vector of floats and returns decoded output.
get_generation_config(self)get_tokenizer(self)set_generation_config(self, config)Attributes
- __annotations__ = {}#
- __class__#
alias of
pybind11_type
- __delattr__(name, /)#
Implement delattr(self, name).
- __dir__()#
Default dir() implementation.
- __eq__(value, /)#
Return self==value.
- __format__(format_spec, /)#
Default object formatter.
Return str(self) if format_spec is empty. Raise TypeError otherwise.
- __ge__(value, /)#
Return self>=value.
- __getattribute__(name, /)#
Return getattr(self, name).
- __getstate__()#
Helper for pickle.
- __gt__(value, /)#
Return self>value.
- __hash__()#
Return hash(self).
- __init__(self: openvino_genai.py_openvino_genai.ASRPipeline, models_path: os.PathLike | str | bytes, device: str, **kwargs) None#
ASRPipeline class constructor. models_path (os.PathLike): Path to the model file. device (str): Device to run the model on (e.g., CPU, GPU).
- __init_subclass__()#
This method is called when a class is subclassed.
The default implementation does nothing. It may be overridden to extend subclasses.
- __le__(value, /)#
Return self<=value.
- __lt__(value, /)#
Return self<value.
- __ne__(value, /)#
Return self!=value.
- __new__(**kwargs)#
- __reduce__()#
Helper for pickle.
- __reduce_ex__(protocol, /)#
Helper for pickle.
- __repr__()#
Return repr(self).
- __setattr__(name, value, /)#
Implement setattr(self, name, value).
- __sizeof__()#
Size of object in memory, in bytes.
- __str__()#
Return str(self).
- __subclasshook__()#
Abstract classes can override this to customize issubclass().
This is invoked early on by abc.ABCMeta.__subclasscheck__(). It should return True, False or NotImplemented. If it returns NotImplemented, the normal algorithm is used. Otherwise, it overrides the normal algorithm (and the outcome is cached).
- _pybind11_conduit_v1_()#
- generate(self: openvino_genai.py_openvino_genai.ASRPipeline, audio_inputs: collections.abc.Sequence[SupportsFloat], generation_config: openvino_genai.py_openvino_genai.ASRGenerationConfig | None = None, streamer: collections.abc.Callable[[str], int | None] | openvino_genai.py_openvino_genai.StreamerBase | None = None, **kwargs) openvino_genai.py_openvino_genai.ASRDecodedResults#
High level generate that receives raw speech as a vector of floats and returns decoded output.
- Parameters:
audio_inputs (list[float]) – inputs in the form of list of floats. Required to be normalized to near [-1, 1] range and have 16k Hz sampling rate.
generation_config (ASRGenerationConfig) – generation_config
streamer – streamer either as a lambda with a boolean returning flag whether generation should be stopped. Streamer supported for short-form audio (< 30 seconds) with return_timestamps=False only
:type : Callable[[str], bool], ov.genai.StreamerBase
- Parameters:
kwargs – arbitrary keyword arguments with keys corresponding to ASRGenerationConfig fields.
:type : dict
- Returns:
return results in decoded form
- Return type:
ASRGenerationConfig
Common parameters:
- Parameters:
Whisper parameters:
- Parameters:
decoder_start_token_id (int) – Corresponds to the “<|startoftranscript|>” token.
pad_token_id (int) – Padding token id.
translate_token_id (int) – Translate token id.
transcribe_token_id (int) – Transcribe token id.
prev_sot_token_id (int) – Corresponds to the “<|startofprev|>” token.
no_timestamps_token_id (int) – No timestamps token id.
begin_suppress_tokens (list[int]) – A list containing tokens that will be suppressed at the beginning of the sampling process.
suppress_tokens (list[int]) – A list containing the non-speech tokens that will be suppressed during generation.
max_initial_timestamp_index (int) – Maximum initial timestamp index.
is_multilingual (bool) – Whether the model is multilingual.
task (Optional[str]) – Task to use for generation, either “translate” or “transcribe”. Can be set for multilingual models only.
lang_to_id (dict[str, int]) – Language token to token_id map. Initialized from the generation_config.json lang_to_id dictionary.
word_timestamps (bool) – If true the pipeline will return word-level timestamps. When enabled word_timestamps=True property should be passed to ASRPipeline constructor: ASRPipeline(“model_path”, “CPU”, word_timestamps=True)
alignment_heads (list[tuple[int, int]]) – Encoder attention alignment heads used for word-level timestamps prediction. Each pair represents (layer_index, head_index).
initial_prompt (Optional[str]) –
Initial prompt tokens passed as a previous transcription (after <|startofprev|> token) to the first processing window. Can be used to steer the model to use particular spellings or styles.
Example:
result = pipeline.generate(raw_speech) # He has gone and gone for good answered Paul Icrom who... result = pipeline.generate(raw_speech, initial_prompt="Polychrome") # He has gone and gone for good answered Polychrome who...
hotwords (Optional[str]) –
Hotwords tokens passed as a previous transcription (after <|startofprev|> token) to all processing windows. Can be used to steer the model to use particular spellings or styles.
Example:
result = pipeline.generate(raw_speech) # He has gone and gone for good answered Paul Icrom who... result = pipeline.generate(raw_speech, hotwords="Polychrome") # He has gone and gone for good answered Polychrome who...
Qwen3-ASR parameters:
- Parameters:
context (Optional[str]) – System prompt context prepended to Qwen3-ASR transcription requests.
For generic generation parameters (max_length, max_new_tokens, num_beams, temperature, etc.) see GenerationConfig documentation.
- get_generation_config(self: openvino_genai.py_openvino_genai.ASRPipeline) openvino_genai.py_openvino_genai.ASRGenerationConfig#
- get_tokenizer(self: openvino_genai.py_openvino_genai.ASRPipeline) openvino_genai.py_openvino_genai.Tokenizer#
- set_generation_config(self: openvino_genai.py_openvino_genai.ASRPipeline, config: openvino_genai.py_openvino_genai.ASRGenerationConfig) None#