python/FunASR-XL.git

			@@ -58,6 +58,22 @@
			#### [RNN-T-online model]()
			Undo

			#### [MFCCA Model](https://www.modelscope.cn/models/NPU-ASLP/speech_mfcca_asr-zh-cn-16k-alimeeting-vocab4950/summary)
			For more model detailes, please refer to [docs](https://www.modelscope.cn/models/NPU-ASLP/speech_mfcca_asr-zh-cn-16k-alimeeting-vocab4950/summary)
			```python
			from modelscope.pipelines import pipeline
			from modelscope.utils.constant import Tasks

			inference_pipeline = pipeline(
			task=Tasks.auto_speech_recognition,
			model='NPU-ASLP/speech_mfcca_asr-zh-cn-16k-alimeeting-vocab4950',
			model_revision='v3.0.0'
			)

			rec_result = inference_pipeline(audio_in='https://isv-data.oss-cn-hangzhou.aliyuncs.com/ics/MaaS/ASR/test_audio/asr_example_zh.wav')
			print(rec_result)
			```

			#### API-reference
			##### Define pipeline
			- `task`: `Tasks.auto_speech_recognition`
			@@ -72,7 +88,7 @@
			- pcm_path, `e.g.`: asr_example.pcm,
			- audio bytes stream, `e.g.`: bytes data from a microphone
			- audio sample point，`e.g.`: `audio, rate = soundfile.read("asr_example_zh.wav")`, the dtype is numpy.ndarray or torch.Tensor
			- wav.scp, kaldi style wav list (`wav_id \t wav_path``), `e.g.`:
			- wav.scp, kaldi style wav list (`wav_id \t wav_path`), `e.g.`:
			```text
			asr_example1 ./audios/asr_example1.wav
			asr_example2 ./audios/asr_example2.wav
			@@ -94,6 +110,8 @@
			- `njob`: only used for CPU inference (`gpu_inference`=`false`), `64` (Default), the number of jobs for CPU decoding
			- `checkpoint_dir`: only used for infer finetuned models, the path dir of finetuned models
			- `checkpoint_name`: only used for infer finetuned models, `valid.cer_ctc.ave.pb` (Default), which checkpoint is used to infer
			- `decoding_mode`: `normal` (Default), decoding mode for UniASR model(fast、normal、offline)
			- `hotword_txt`: `None` (Default), hotword file for contextual paraformer model(the hotword file name ends with .txt")

			- Decode with multi GPUs:
			```shell
			@@ -176,6 +194,21 @@
			- `max_epoch`: number of training epoch
			- `lr`: learning rate

			- Training data formats：
			```sh
			cat ./example_data/text
			BAC009S0002W0122 而对楼市成交抑制作用最大的限购
			BAC009S0002W0123 也成为地方政府的眼中钉
			english_example_1 hello world
			english_example_2 go swim 去游泳

			cat ./example_data/wav.scp
			BAC009S0002W0122 /mnt/data/wav/train/S0002/BAC009S0002W0122.wav
			BAC009S0002W0123 /mnt/data/wav/train/S0002/BAC009S0002W0123.wav
			english_example_1 /mnt/data/wav/train/S0002/english_example_1.wav
			english_example_2 /mnt/data/wav/train/S0002/english_example_2.wav
			```

			- Then you can run the pipeline to finetune with:
			```shell
			python finetune.py
			@@ -210,4 +243,4 @@
			--njob 64 \
			--checkpoint_dir "./checkpoint" \
			--checkpoint_name "valid.cer_ctc.ave.pb"
			```
			```