python/FunASR-XL.git

			@@ -1,7 +1,9 @@
			([简体中文](./README_zh.md)\|English)

			# Punctuation Restoration

			> Note:
			> The modelscope pipeline supports all the models in [model zoo](https://alibaba-damo-academy.github.io/FunASR/en/modelscope_models.html#pretrained-models-on-modelscope) to inference and finetune. Here we take the model of the punctuation model of CT-Transformer as example to demonstrate the usage.
			> The modelscope pipeline supports all the models in [model zoo](https://alibaba-damo-academy.github.io/FunASR/en/model_zoo/modelscope_models.html#pretrained-models-on-modelscope) to inference and finetune. Here we take the model of the punctuation model of CT-Transformer as example to demonstrate the usage.

			## Inference

			@@ -55,7 +57,7 @@
			### API-reference
			#### Define pipeline
			- `task`: `Tasks.punctuation`
			- `model`: model name in [model zoo](https://alibaba-damo-academy.github.io/FunASR/en/modelscope_models.html#pretrained-models-on-modelscope), or model path in local disk
			- `model`: model name in [model zoo](https://alibaba-damo-academy.github.io/FunASR/en/model_zoo/modelscope_models.html#pretrained-models-on-modelscope), or model path in local disk
			- `ngpu`: `1` (Default), decoding on GPU. If ngpu=0, decoding on CPU
			- `output_dir`: `None` (Default), the output path of results if set
			- `model_revision`: `None` (Default), setting the model version
			@@ -70,17 +72,17 @@
			### Inference with multi-thread CPUs or multi GPUs
			FunASR also offer recipes [egs_modelscope/punctuation/TEMPLATE/infer.sh](https://github.com/alibaba-damo-academy/FunASR/blob/main/egs_modelscope/punctuation/TEMPLATE/infer.sh) to decode with multi-thread CPUs, or multi GPUs. It is an offline recipe and only support offline model.

			- Setting parameters in `infer.sh`
			- `model`: model name in [model zoo](https://alibaba-damo-academy.github.io/FunASR/en/modelscope_models.html#pretrained-models-on-modelscope), or model path in local disk
			- `data_dir`: the dataset dir needs to include `punc.txt`
			- `output_dir`: output dir of the recognition results
			- `gpu_inference`: `true` (Default), whether to perform gpu decoding, set false for CPU inference
			- `gpuid_list`: `0,1` (Default), which gpu_ids are used to infer
			- `njob`: only used for CPU inference (`gpu_inference`=`false`), `64` (Default), the number of jobs for CPU decoding
			- `checkpoint_dir`: only used for infer finetuned models, the path dir of finetuned models
			- `checkpoint_name`: only used for infer finetuned models, `punc.pb` (Default), which checkpoint is used to infer
			#### Settings of `infer.sh`
			- `model`: model name in [model zoo](https://alibaba-damo-academy.github.io/FunASR/en/model_zoo/modelscope_models.html#pretrained-models-on-modelscope), or model path in local disk
			- `data_dir`: the dataset dir needs to include `punc.txt`
			- `output_dir`: output dir of the recognition results
			- `gpu_inference`: `true` (Default), whether to perform gpu decoding, set false for CPU inference
			- `gpuid_list`: `0,1` (Default), which gpu_ids are used to infer
			- `njob`: only used for CPU inference (`gpu_inference`=`false`), `64` (Default), the number of jobs for CPU decoding
			- `checkpoint_dir`: only used for infer finetuned models, the path dir of finetuned models
			- `checkpoint_name`: only used for infer finetuned models, `punc.pb` (Default), which checkpoint is used to infer

			- Decode with multi GPUs:
			#### Decode with multi GPUs:
			```shell
			bash infer.sh \
			--model "damo/punc_ct-transformer_zh-cn-common-vocab272727-pytorch" \
			@@ -90,7 +92,7 @@
			--gpu_inference true \
			--gpuid_list "0,1"
			```
			- Decode with multi-thread CPUs:
			#### Decode with multi-thread CPUs:
			```shell
			bash infer.sh \
			--model "damo/punc_ct-transformer_zh-cn-common-vocab272727-pytorch" \
			@@ -99,7 +101,6 @@
			--gpu_inference false \
			--njob 1
			```


			## Finetune with pipeline