Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Panhapich
/
khmer-sp-8k
like
1
Khmer
English
khmer
cambodia
sentencepiece
tokenizer
unigram
code-switching
License:
other
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
main
khmer-sp-8k
696 kB
Ctrl+K
Ctrl+K
1 contributor
History:
19 commits
Panhapich
Upload khmer_segmentation.py with huggingface_hub
ce3b16e
verified
28 days ago
.gitattributes
Safe
1.52 kB
initial commit
30 days ago
README.md
Safe
4.03 kB
Upload README.md with huggingface_hub
28 days ago
USAGE.md
Safe
5.52 kB
Rename USAGE to USAGE.md
29 days ago
gazetteer.json
Safe
268 Bytes
Upload gazetteer.json with huggingface_hub
28 days ago
khmer_segmentation.py
8.71 kB
Upload khmer_segmentation.py with huggingface_hub
28 days ago
khmer_segmentation_sentencepiece_approach.md
Safe
15.8 kB
Upload khmer_segmentation_sentencepiece_approach.md with huggingface_hub
28 days ago
khmer_sp.model
Safe
442 kB
xet
Upload khmer_sp.model with huggingface_hub
28 days ago
khmer_sp.vocab
217 kB
Upload khmer_sp.vocab with huggingface_hub
28 days ago
latin_exceptions.json
Safe
210 Bytes
Upload latin_exceptions.json with huggingface_hub
28 days ago
tokenizer_info.json
1.31 kB
Upload tokenizer_info.json with huggingface_hub
28 days ago