gpt4all/.codespellrc at d59c77ac553cbc51776811a3fee0229a06cafd45 - gpt4all - gitea: Gitea Service

nomic-ai/gpt4all

mirror of https://github.com/nomic-ai/gpt4all.git synced 2024-09-20 09:37:39 +03:00

Aaron Miller bbcee1ced5 New tokenizer implementation for MPT and GPT-J

Improves output quality by making these tokenizers more closely
match the behavior of the huggingface `tokenizers` based BPE
tokenizers these models were trained with.

Featuring:
 * Fixed unicode handling (via ICU)
 * Fixed BPE token merge handling
 * Complete added vocabulary handling

2023-05-30 12:05:57 -04:00

5 lines

81 B

Plaintext

Raw Blame History

 [codespell]
 skip = .git,*.pdf,*.svg,*_tokenizer_config.h
 #
 # ignore-words-list =