Commit Graph

75 Commits

Author SHA1 Message Date
Traubert
1e349f269c Fix type 2020-09-24 12:15:08 +03:00
Joerg Tiedemann
f9a44bdb99 merged 2020-09-18 23:08:56 +03:00
Joerg Tiedemann
87a5354de5 changes to tatoeba recipes 2020-09-18 23:05:46 +03:00
Jörg Tiedemann
a61cf48443 add option to skip sentence piecce vocabs but use marian_vocab instead 2020-09-16 19:33:19 +03:00
tiedemann
913d31472e
Merge pull request #25 from Helsinki-NLP/sam-fixes
Fix typo
2020-09-16 09:28:46 +03:00
Joerg Tiedemann
c564bd1f56 fix in fetching data for Sami languages 2020-09-16 09:25:36 +03:00
Traubert
1da54c4155 Fix typo 2020-09-15 15:02:20 +03:00
Jörg Tiedemann
58fbf0bdd8 back to old subword model names 2020-09-14 08:53:57 +03:00
Jörg Tiedemann
c2798e9758 plain text vocab files from spm models 2020-09-13 22:17:21 +03:00
Jörg Tiedemann
24e92de56a proper release packages for models with internal sentence piece vocabs 2020-09-13 00:00:15 +03:00
Jörg Tiedemann
666b2b8462 internal sentence piece models in transformers 2020-09-12 16:16:01 +03:00
Jörg Tiedemann
ddafb43d66 removed dependence on moses tools in preprocessing script for released spm packages 2020-09-12 14:42:10 +03:00
Jörg Tiedemann
c0cb356417 added acknowledgements 2020-09-12 12:01:02 +03:00
Jörg Tiedemann
16eef8e45d moved project makefiles to lib/projects 2020-09-10 12:12:44 +03:00
Jörg Tiedemann
1a6e29275d dev data is now uniq to avoid overlaps with test data 2020-09-09 23:21:07 +03:00
Jörg Tiedemann
3735af4ec1 more documentation 2020-09-07 23:00:01 +03:00
Jörg Tiedemann
a47c292152 pivoting and documentation 2020-09-05 22:19:00 +03:00
Jörg Tiedemann
ad828c3124 started tutorial and fixes to backtranslate makefile 2020-09-05 00:16:22 +03:00
Tiedemann Jörg
d11f74ce41 added bpe submodule 2020-09-04 15:34:20 +03:00
Tiedemann
96eaad2d05 added possibility to fetch moses file from ObjectStore (instead of reading with opus_read) 2020-09-03 22:04:44 +03:00
Joerg Tiedemann
971ece9606 fix tatoeba data labels 2020-09-03 07:55:44 +03:00
Tiedemann
1435b7849a moved allas recipes to a different makefile 2020-09-02 16:35:35 +03:00
Tiedemann
2332732577 make compatible with mac osx and include submodules for required tools 2020-09-02 15:52:34 +03:00
Joerg Tiedemann
639bd2adda started documentation of project specific models 2020-08-28 15:51:37 +03:00
Joerg Tiedemann
2c04e48dbe fixed an important bug in data merging 2020-08-28 11:52:46 +03:00
Joerg Tiedemann
e31550a3ad enabled fetching OPUS data instead of reading local files if necessary 2020-08-28 10:53:11 +03:00
Joerg Tiedemann
94eeec13eb take away dependence on local OPUS files for finding data 2020-08-27 22:36:50 +03:00
Joerg Tiedemann
831ee89f76 fixed bug in env.mk 2020-08-26 22:18:12 +03:00
Joerg Tiedemann
596dd993a5 more documentation 2020-08-26 21:45:03 +03:00
Joerg Tiedemann
a8b54f5311 some info about training added 2020-08-26 15:12:38 +03:00
Joerg Tiedemann
2f8a37cc92 more details about data compilation added 2020-08-26 14:31:50 +03:00
Joerg Tiedemann
dac6070069 started some more documentation 2020-08-26 09:59:24 +03:00
Joerg Tiedemann
f2a413b740 minor cleanup in env 2020-08-26 01:01:44 +03:00
Joerg Tiedemann
4c35456038 cleanup in data makefile 2020-08-26 00:44:02 +03:00
Joerg Tiedemann
9375f37886 missing makefile added 2020-08-25 22:42:33 +03:00
Joerg Tiedemann
0e27198048 store and fetch work data 2020-08-22 23:51:37 +03:00
Joerg Tiedemann
d7252e32b7 tatoeba monolingual data 2020-08-05 00:00:24 +03:00
Joerg Tiedemann
6bf0207cc6 list of models added 2020-08-03 11:58:51 +03:00
Joerg Tiedemann
c9fcb7f35d tatoeba langgroup models 2020-08-02 11:38:42 +03:00
Joerg Tiedemann
1b913277b3 tatoeba language group models with various sample sizews 2020-07-25 22:52:33 +03:00
Joerg Tiedemann
5493aeddb4 fixed a problem with lang group targets 2020-07-14 21:40:49 +03:00
Joerg Tiedemann
068b82cc1d fixed a bug in eval/dist groups 2020-07-14 13:29:06 +03:00
Joerg Tiedemann
e2edc4195a result tables for language groups and minor fixes for start scripts in Tatoeba challenge 2020-07-10 11:59:37 +03:00
Joerg Tiedemann
ec6d7c7142 tatoeba langgroups 2020-07-04 23:37:39 +03:00
Joerg Tiedemann
7df91a9eaa language group jobs with some more documentation 2020-06-29 12:26:45 +03:00
Joerg Tiedemann
62c9414122 lang groups 2020-06-29 00:15:35 +03:00
Joerg Tiedemann
46a0b2b15a fixed dist-packaging 2020-06-27 13:56:51 +03:00
Joerg Tiedemann
e2bc2acb3b re-organised targets for multilingual models of language groups 2020-06-27 12:29:50 +03:00
Joerg Tiedemann
9e186d82d6 bugfix in tatoeba data extraction for multilingual data files (language code clash) 2020-06-25 00:45:25 +03:00
Joerg Tiedemann
844f8bf72a removed unnecessary pre-processing for chinese 2020-06-19 16:12:06 +03:00