Nominatim

mirror of https://github.com/osm-search/Nominatim.git synced 2024-11-25 19:35:02 +03:00

Author	SHA1	Message	Date
Tareq Al-Ahdal	ac467c7a2d	Enhanced the implementation of OSM views GeoTIFF import functionality	2022-10-01 11:01:49 +02:00
Tareq Al-Ahdal	c85b74497b	Initial implementation of GeoTIFF import functionality	2022-10-01 11:01:49 +02:00
Sarah Hoffmann	f4d3ae6f70	consolidate indexes over geometry_sectors The index over geometry_sectors are mainly used for ordering the places which need indexing. That means they function effectively as a TODO list. Consolodate them so that they always only contain the places which are still to do. Also add the appropriate index for the boundary indexing phase.	2022-09-21 10:38:58 +02:00
Sarah Hoffmann	dddfa3a075	ignore irrelevant extra tags on address interpolations When deciding if an address interpolation has address information, only look for addr:street and addr:place. If they are not there go looking for the address on the address nodes. Ignores irrelevant tags like addr:inclusion. Fixes #2797.	2022-08-13 14:07:06 +02:00
Sarah Hoffmann	487e81fe3c	more invalidations when boundary changes rank When a boundary or place changes its address rank, all places where it participates as address need to be potentially reindexed. Also use the computed rank when testing place nodes against boundaries. Boundaries are computed earlier. Fixes #2794.	2022-08-12 09:48:46 +02:00
Sarah Hoffmann	51b6d16dc6	overhaul the token analysis interface The functional split betweenthe two functions is now that the first one creates the ID that is used in the word table and the second one creates the variants. There no longer is a requirement that the ID is the normalized version. We might later reintroduce the requirement that a normalized version be available but it doesn't necessarily need to be through the ID. The function that creates the ID now gets the full PlaceName. That way it might take into account attributes that were set by the sanitizers. Finally rename both functions to something more sane.	2022-07-29 15:14:11 +02:00
Sarah Hoffmann	c8873d34af	harmonize interface of token analysis module The configure() function now receives a Transliterator object instead of the ICU rules. This harmonizes the parameters with the create function.	2022-07-29 10:43:07 +02:00
Sarah Hoffmann	6d41046b15	add support for external sanitizer modules	2022-07-25 16:10:19 +02:00
Sarah Hoffmann	7b7203c149	add function for loading plugin modules Loads modules for configurable code like tokenizers, sanitizers, etc. Supports internal modules, external libraries and code from the project directory.	2022-07-25 16:10:10 +02:00
Sarah Hoffmann	cd4bcea894	ignore API parameters in array notation PHP automatically parses parameters in an array notation(foo[]) into array types. Ignore these parameters as 'unknown'. Fixes #2763.	2022-07-23 10:51:44 +02:00
Kian-Meng Ang	f5e52e748f	docs: fix typos	2022-07-20 22:05:31 +08:00
Sarah Hoffmann	9963261d8d	add type annotations to special phrase importer	2022-07-18 09:54:29 +02:00
Sarah Hoffmann	62eedbb8f6	add type hints for sanitizers	2022-07-18 09:47:57 +02:00
Sarah Hoffmann	fc254fc744	adapt use of Connection in bdd tests to name change	2022-07-18 09:47:57 +02:00
Sarah Hoffmann	aaf2b6032e	fix uses of config.get_path() to expect None	2022-07-18 09:47:57 +02:00
Sarah Hoffmann	b1903f0fbf	Merge pull request #2761 from lonvia/repair-index-analysis Repair `admin --analyse-indexing`	2022-07-18 09:38:08 +02:00
marc tobias	c70ca7f57b	In tests for PHP 8 disable Just-in-time, it conflicts with tools that determine coverage	2022-07-09 22:03:48 +02:00
Sarah Hoffmann	4b12d52ef5	convert admin --analyse-indexing to new indexing method A proper run of indexing requires the place information from the analyzer. Add the pre-processing of place data, so the right information is handed into the update function.	2022-07-07 16:20:08 +02:00
Sarah Hoffmann	cbbcbb1fd7	move country_info into data submodule	2022-07-06 11:08:36 +02:00
Sarah Hoffmann	bce93d60bd	move PlaceInfo into data submodule This data structure is shared between indexer and tokenizer.	2022-07-06 10:54:47 +02:00
Sarah Hoffmann	69e51aebab	test: avoid column names with upper-case letters This may cause problems when the column names get quoted.	2022-07-05 09:12:55 +02:00
Marc Tobias	ccf119206d	PHP 8 behaves slightly different with in_array and usort	2022-07-03 10:55:34 +02:00
Sarah Hoffmann	3dd7410bb7	bdd: correctly skip postcode tests for legacy	2022-06-23 23:42:31 +02:00
Sarah Hoffmann	93d5be097a	bdd: do not expect legacy word table to be without empty tokens It can happen for bogus names and this will not get fixed anymore.	2022-06-23 23:42:31 +02:00
Sarah Hoffmann	6eb9044353	adapt search algorithm to new postcode format in word	2022-06-23 23:42:31 +02:00
Sarah Hoffmann	612d34930b	handle postcodes properly on word table updates update_postcodes_from_db() needs to do the full postcode treatment in order to derive the correct word table entries.	2022-06-23 23:42:31 +02:00
Sarah Hoffmann	0f00f4968c	fix up BDD tests for postcode changes Includes smaller code fixes found by the tests.	2022-06-23 23:42:31 +02:00
Sarah Hoffmann	7b6ec4fc6c	add tests for discarding bad postcodes	2022-06-23 23:42:31 +02:00
Sarah Hoffmann	80ea13437d	move postcode matcher in a separate file	2022-06-23 23:42:31 +02:00
Sarah Hoffmann	4885fdf0f9	add class for online centroid computation	2022-06-23 23:42:31 +02:00
Sarah Hoffmann	18864afa8a	postcodes: introduce a default pattern for countries without postcodes	2022-06-23 23:42:31 +02:00
Sarah Hoffmann	9172696324	postcodes: add support for optional spaces	2022-06-23 23:42:31 +02:00
Sarah Hoffmann	baee6f3de0	postcodes: strip leading country codes	2022-06-23 23:42:31 +02:00
Sarah Hoffmann	28ab2f6048	add postcodes patterns without optional spaces	2022-06-23 23:42:31 +02:00
Sarah Hoffmann	90d4d339db	initial postcode cleaner for simple patterns Moves postcodes that are either in countries without a postcode system or don't correspond to the local pattern for postcodes into a field for a normal address part. Makes them searchable but not as a special address. This has two consequences: they are no longer a skippable part of the address and the postcodes cannot be searched on their own.	2022-06-23 23:42:31 +02:00
Sarah Hoffmann	8080625747	remove postcodes from countries that don't have them The postcodes will only be removed as a 'computed postcode' they are still searchable for the given object.	2022-06-23 23:42:31 +02:00
Sarah Hoffmann	d8623d6818	bdd: remove support for scenes Only keep support for the special point geometry 'country:xx'.	2022-06-17 11:54:18 +02:00
Sarah Hoffmann	6c58a4c46c	bdd: move query tests from scene to grid description	2022-06-17 11:54:18 +02:00
Sarah Hoffmann	19f67e167c	bdd: remove step for scene setup	2022-06-17 11:54:18 +02:00
Sarah Hoffmann	00d8df6fc3	bdd: move update tests from scenes to grid descriptions	2022-06-17 11:54:18 +02:00
Sarah Hoffmann	02068aec7f	bdd: move import tests from scenes to grid descriptions	2022-06-17 11:54:18 +02:00
Sarah Hoffmann	3493d317e4	bdd: clear lof buffer after a successful import run	2022-06-17 11:54:18 +02:00
Sarah Hoffmann	a2b486a5b0	bdd: allow to set an origin of the grid	2022-06-17 11:54:18 +02:00
Sarah Hoffmann	df0142678a	improve address ordering with mixes of place and admin areas Resolves a couple of situations where a mixed use of places areas and administrative boundaries would result in a hierarchy that did not properly respect the contains relation.	2022-06-16 10:44:16 +02:00
Sarah Hoffmann	15cf7dd416	add testcase for #2551 This test proves that places that are linked need to be reindexed.	2022-06-05 21:39:17 +02:00
Sarah Hoffmann	cbb4749996	change indexing order for interpolations Interpolations are now indexed after rank 30 objects. The housenumber nodes no longer need information from the interpolations while the interpolations can make use of precomputed postcodes.	2022-06-02 15:16:46 +02:00
Sarah Hoffmann	8a0e3e2f3d	Merge pull request #2732 from lonvia/fix-ordering-address-parts Fix order when searching for addr:* components	2022-05-31 20:26:05 +02:00
Sarah Hoffmann	bd0e157b91	fix order when searching for addr:* components When matching addr:* components the preference was given to matches that do not intersect with the place.	2022-05-31 16:57:37 +02:00
Sarah Hoffmann	46689df668	custom comparison for SpecialPhrase Duplicate elemination only works when a custom hash/equal function is implemented that is based on the members.	2022-05-30 16:30:41 +02:00
Sarah Hoffmann	e828d0d3f7	move quoting hack to wiki loader The bad quotes around the type for special phrases specifically occure in the Wiki pages, so it should be removed by the loader and not in the generic SpecialPhrase object.	2022-05-30 14:40:33 +02:00
Sarah Hoffmann	cce0e5ea38	convert special phrase loaders to generators Generators simplify the code quite a bit compared to the previous Iterator approach.	2022-05-30 14:12:46 +02:00
Sarah Hoffmann	042e314589	remove the language parameter in the SPWikiLoader Languages must always be configured through config or environment. Also use monkeypatched environment in tests.	2022-05-30 10:26:20 +02:00
Sarah Hoffmann	61d813bfef	add get_str_list() for config Converts a config value written as a comma-sparated list into a Python list of strings.	2022-05-29 13:53:50 +02:00
Sarah Hoffmann	1d203fdb3c	fix bug with keeping linking on updates When moving the finding of linked places to the precomputation stage, it was also moved before the statement where the linked_place_id was removed from the linkee. The result was that the current linkee was excluded when looking for a linked place on updates because it was still linked to the boundary to be updated. Fixed by allowing to either keep the linkage or change to an unlinked place.	2022-05-23 10:55:10 +02:00
Sarah Hoffmann	f314abcfe1	bdd: restrict imports to four languages This mainly restricts the number of country names that are loaded.	2022-05-11 16:40:53 +02:00
Sarah Hoffmann	e74e577029	bdd: recreate functions on template DB Avoids calling function refresh on every scenario. The content won't change between runs.	2022-05-11 15:50:22 +02:00
Sarah Hoffmann	aa0ae610c6	avoid calling OSM servers during bdd tests	2022-05-11 15:33:01 +02:00
Sarah Hoffmann	5ff35d9984	Merge pull request #2707 from lonvia/make-icu-tokenizer-the-default Make ICU tokenizer the default	2022-05-11 08:52:49 +02:00
Marc Tobias	99fa23040a	PHPUnit 9 changed configuration schema slightly	2022-05-10 15:20:43 +02:00
Sarah Hoffmann	adeebec32a	switch tests to ICU tokenizer as default	2022-05-10 14:54:50 +02:00
Sarah Hoffmann	ed6fda6968	Merge pull request #2702 from lonvia/move-country-names-into-includes Clean up country name settings	2022-05-10 09:21:16 +02:00
Marc Tobias	821dabb138	add git commit hash to --version output	2022-05-09 23:56:13 +02:00
Sarah Hoffmann	9d468f6da0	support arbitrary prefixes in country name list This means we can now get rid of the last special cases for names.	2022-05-09 11:55:26 +02:00
Marc Tobias	0de83c4a51	fix typos of name Nominatim	2022-05-05 01:04:47 +02:00
Marc Tobias	a79ab41782	new nominatim --version CLI argument	2022-05-04 01:33:25 +02:00
Sarah Hoffmann	372874e89a	accept any OSM type in street member of associatedStreet This is needed for pedestrian areas mapped as multipolygons and consequently as relations. The lookup in placex guarantees that the referenced OSM object is indeed a street. Fixes #2669.	2022-05-02 09:48:51 +02:00
Sarah Hoffmann	3c68b12176	keep inherited address parts after indexing The inherited housenumber is needed for display output. We can't take the one from the housenumber field because it is already normalized. Remove the inherited address only when reindexing. Fixes #2683.	2022-04-28 21:38:00 +02:00
Sarah Hoffmann	4f59644cc2	add tests for new data invalidation functions	2022-04-14 14:52:13 +02:00
Artem Ziablytskyi	d1479072ae	fix bdd tests and docs	2022-04-07 16:37:51 +02:00
Artem Ziablytskyi	9a56e53d50	use ISO3166-2-lvl<admin_level> instead of typeLabel prefix	2022-04-07 16:37:51 +02:00
Artem Ziablytskyi	6bee188f24	Change the key to `<addresspart_type>-ISO3166-2` to support xml response correctly	2022-04-07 16:37:51 +02:00
Artem Ziablytskyi	82dbcbb12a	add `<addresspart_type>:ISO3166-2` field to response address details	2022-04-07 16:37:51 +02:00
Artem Ziablytskyi	76c146f326	add `state_code` field to response address details	2022-04-07 16:37:51 +02:00
Sarah Hoffmann	fd4ab3f262	Merge pull request #2629 from tareqpi/country-names-yaml-configuration Move default country names into yaml configuration	2022-04-04 09:04:25 +02:00
Tareq Al-Ahdal	e9f979b67b	'read_config' is no longer a fixture add 'read_config' to test cases that need it	2022-04-01 22:52:17 +08:00
Tareq Al-Ahdal	a323b8f63a	test for loading special characters from country_settings.yaml	2022-04-01 21:58:57 +08:00
Tareq Al-Ahdal	9411c14fd2	fix reset country info before loading custom data	2022-04-01 21:55:34 +08:00
Tareq Al-Ahdal	8525e7542f	custom country config loads correctly	2022-04-01 21:46:56 +08:00
Sarah Hoffmann	de18cd1523	add test for new table_has_column function	2022-03-31 15:55:20 +02:00
Sarah Hoffmann	36a1560117	add migration to mark internal country names	2022-03-31 15:55:20 +02:00
Tareq Al-Ahdal	b5f311d6bc	separate unit test function into three functions	2022-03-30 22:06:59 +08:00
Tareq Al-Ahdal	9db13aac72	Added unit tests for loading country info from yaml file	2022-03-25 22:22:44 +08:00
Sarah Hoffmann	a0ed80d821	restore the tokenizer directory when missing Automatically repopulate the tokenizer/ directory with the PHP stub and the postgresql module, when the directory is missing. This allows to switch working directories and in particular run the service from a different maschine then where it was installed. Users still need to make sure that .env files are set up correctly or they will shoot themselves in the foot. See #2515.	2022-03-20 11:31:42 +01:00
Tareq Al-Ahdal	943e5fe699	Revert the removal of new line at the end of the file	2022-03-18 06:07:48 +08:00
Tareq Al-Ahdal	83b4b8d9c1	reattach 'name:' prefix to keys	2022-03-18 05:46:23 +08:00
Tareq Al-Ahdal	d0c1b73fb3	remove duplicate values	2022-03-18 02:43:42 +08:00
Tareq Al-Ahdal	6be2077d92	Merge branch 'master' into country-names-yaml-configuration	2022-03-18 02:36:12 +08:00
Tareq Al-Ahdal	456d439e97	Reformatting of country keys	2022-03-18 02:23:11 +08:00
Sarah Hoffmann	23de4c7aca	adapt ParameterParser tests to new key list	2022-03-17 11:45:05 +01:00
Sarah Hoffmann	e133476c35	merge linked names correctly into namedetails Convert the '_place_' entries back to normal entries before returning them in the 'namedetails' section. If the name field is duplicated, kept the '_place_' notation. This preserves the previous behaviour before _place_ names were introduces but adds the additional names from the linked place for reference.	2022-03-17 11:02:02 +01:00
Sarah Hoffmann	524dc64ab7	make sure outputs take into account linked place names	2022-03-16 21:44:52 +01:00
Sarah Hoffmann	42cd021d04	save differing linked polace names in extra fields This keeps the names tracable and ensures that all names are searchable when they differ. Do not keep names when they are exactly the same to save some space. Linked names are cleaned out before relinking.	2022-03-16 16:38:52 +01:00
Sarah Hoffmann	ef98a85b05	correctly handle single-point interpolations in reverse Lookup in location_property_osmline needs to be special cased for startnumber = endnumber. Also adds tests for the case. Fixes #2680.	2022-03-16 11:19:09 +01:00
Sarah Hoffmann	0a9f971e44	add tests for new analyzed housenumbers	2022-03-01 09:34:32 +01:00
Sarah Hoffmann	89e1446131	bdd: disable some housenumber tests for legacy Optional spaces in housenumbers are not supported by legacy tokenizer, so disable those tests.	2022-03-01 09:34:32 +01:00
Sarah Hoffmann	f03a05f6bb	add new analyser for houenumbers This analyser makes spaces optional.	2022-03-01 09:34:32 +01:00
Sarah Hoffmann	837d44391c	move generation of normalized token form to analyzer This gives the analyzer more flexibility in choosing the normalized form. In particular, an analyzer creating different variants can choose the variant that will be used as the canonical form.	2022-03-01 09:34:32 +01:00
Sarah Hoffmann	1d82569f6d	add tests for country updates	2022-02-24 16:18:49 +01:00
Sarah Hoffmann	f74228830d	bdd: run full import on tests This uncovered a couple of outdated/wrong tests which have been fixed, too.	2022-02-24 14:27:51 +01:00
Sarah Hoffmann	0e11ca9b76	add test that interpolations are found by odd/even	2022-02-10 11:23:51 +01:00
Sarah Hoffmann	a6b4e8ff67	add tests for housenumber-as-name feature	2022-02-07 11:45:12 +01:00
Sarah Hoffmann	38c3ef3da0	add tests for get_string_list() Renaming test file for sanitizer config because pytest requires unique names for test files.	2022-02-07 11:22:24 +01:00
Sarah Hoffmann	610f2cc254	sanitizer: move helpers into a configuration class	2022-02-07 10:48:00 +01:00
Sarah Hoffmann	a79a3210e6	implement is-a-name option for housenumbers	2022-02-07 09:27:11 +01:00
Sarah Hoffmann	b6fa121f53	remove tests for closest housenumber function	2022-01-27 16:21:45 +01:00
Sarah Hoffmann	64abc90d30	use new tiger step column for queries	2022-01-27 14:08:08 +01:00
Sarah Hoffmann	6b89624f33	adapt frontend to new interpolation table layout	2022-01-27 11:14:55 +01:00
Sarah Hoffmann	4b28b4fed4	adapt BDD tests for new interpolation style	2022-01-27 11:14:55 +01:00
Sarah Hoffmann	c170d323d9	add tests for cleaning housenumbers	2022-01-20 23:47:20 +01:00
Sarah Hoffmann	d09db09849	adapt ICU tets to new housenumber sanitizer Restrict tests to making sure that handing in multiple housenumbers works.	2022-01-20 16:05:49 +01:00
Sarah Hoffmann	3741afa6dc	generalize filter-kind parameter for sanatizers Now behaves the same for tag_analyzer_by_language and clean_housenumbers. Adds tests.	2022-01-20 15:42:42 +01:00
Sarah Hoffmann	560a006892	add pytest config We are using custom marks now which need to be registered to avoid warnings.	2022-01-20 15:38:02 +01:00
Sarah Hoffmann	4774e45218	clean_housenumbers: make kinds and delimiters configurable Also adds unit tests for various options.	2022-01-20 12:07:12 +01:00
Sarah Hoffmann	206ee87188	factor out housenumber splitting into sanitizer	2022-01-19 17:27:50 +01:00
Sarah Hoffmann	b453b0ea95	introduce mutation variants to generic token analyser Mutations are regular-expression-based replacements that are applied after variants have been computed. They are meant to be used for variations on character level. Add spelling variations for German umlauts.	2022-01-18 11:09:21 +01:00
Sarah Hoffmann	c3788d765e	add consistent SPDX copyright headers	2022-01-03 16:23:58 +01:00
Sarah Hoffmann	ab6f35d83a	Merge pull request #2553 from lonvia/revert-street-matching-to-full-names Revert street matching to full names	2021-12-14 15:52:34 +01:00
Sarah Hoffmann	f9b56a8581	correctly match abbreviated addr:street This only works when addr:street is abbreviated and the street name isn't. It does not work the other way around.	2021-12-08 21:58:43 +01:00
Sarah Hoffmann	04857d32cd	enable PHPUnit 9 for coverage A couple of functions have been renamed.	2021-12-07 12:07:17 +01:00
Sarah Hoffmann	109cdce92c	php unit: replace deprecated regex assert The regEx assertion has been renamed in PHPUnit 9.5 and causes deprecation warnings.	2021-12-07 11:34:21 +01:00
Sarah Hoffmann	b7554d9ed8	php unit: don't enforce a name on the test database Also gets rid of a PHPUnit deprecation warning.	2021-12-07 11:31:45 +01:00
Sarah Hoffmann	6106f1a32e	php test: class must be called like the file	2021-12-07 11:20:38 +01:00
Sarah Hoffmann	7f7d2fd5b3	skip most addr: tags with suffixes Only one addr: tag can be processed currently, so make sure it is the one without suffixes to not get odd data. addr:street is the exception because it uses a different matching mechanism.	2021-12-06 14:55:10 +01:00
Sarah Hoffmann	5e435b41ba	ICU: matching any street name will do again	2021-12-06 14:26:08 +01:00
Sarah Hoffmann	44cfce1ca4	revert to using full names for street name matching Using partial names turned out to not work well because there are often similarly named streets next to each other. It also prevents us from being able to take into account all addr:street:* tags. This change gets all the full term tokens for the addr:street tags from the DB. As they are used for matching only, we can assume that the term must already be there or there will be no match. This avoid creating unused full name tags.	2021-12-06 11:38:38 +01:00
Sarah Hoffmann	5a9fb6eaf7	specify text type in test SQL Older version of postgres fail otherwise.	2021-12-03 13:56:23 +01:00
Sarah Hoffmann	54d35ddfe9	split cli tests by subcommand and extend coverage	2021-12-02 23:45:48 +01:00
Sarah Hoffmann	14a78f55cd	more unit tests for tokenizers	2021-12-02 15:46:36 +01:00
Sarah Hoffmann	7617a9316e	extend API unit tests	2021-12-01 20:48:29 +01:00
Sarah Hoffmann	a52ed366e4	add tests for migration	2021-12-01 20:27:40 +01:00
Sarah Hoffmann	7be164e2a5	more testing for refresh functions	2021-12-01 14:58:54 +01:00
Sarah Hoffmann	a24f25c0d8	more tests for exec utilities	2021-12-01 14:23:51 +01:00
Sarah Hoffmann	993b238a41	add more tests for database import	2021-12-01 11:54:58 +01:00
Sarah Hoffmann	bbbfc8201c	add tests for adding additional data Also adds checks that parameters for osm2pgsql are set as expected.	2021-12-01 11:22:46 +01:00
Sarah Hoffmann	6f03a4d6ce	add tests for flatten_config_file and other than yaml formats	2021-12-01 10:24:11 +01:00
Sarah Hoffmann	c8958a22d2	tests: add fixture for making test project directory	2021-11-30 18:01:46 +01:00
Sarah Hoffmann	37afa2180b	generalize fixtures for cli tests	2021-11-30 14:07:39 +01:00
Sarah Hoffmann	b2df8e478a	python test: move single-use fixtures to subdirectories	2021-11-30 12:03:16 +01:00
Sarah Hoffmann	50fccb52be	remove unused test files	2021-11-30 11:44:10 +01:00
Sarah Hoffmann	b90e719da5	organise python tests in subdirectories The directories follow the same structure as the modules in nominatim/.	2021-11-30 11:22:26 +01:00
Sarah Hoffmann	80e0a3cce4	change default rank for highway objects to 30 The highway key is being used more and more for non-ways these days. This clashes with Nominatim's assumption that essentially everything that has a highway tag can be used as the street part of the address. Change the default rank of highway objects to 30 to avoid this. Only the known values for streets keep the rank 26 and are now listed explicitly.	2021-11-24 22:10:40 +01:00
Sarah Hoffmann	10e979e841	only instantiate indexer once for replication Also makes sure that indexer object exists everywhere were needed. See #2518.	2021-11-19 14:48:58 +01:00
Sarah Hoffmann	345c812e43	better error reporting when API script does not exist Check if the API script exists on the expected location before running php-cli. This way we can add a useful hint about the project directory. Fixes #2513.	2021-11-10 11:58:20 +01:00
Sarah Hoffmann	37eeccbf4c	ICU: use normalization from config in PHP The TERM_NORMALIZATION config option is no longer applicable. That was already documented but not yet implemented.	2021-10-27 11:32:44 +02:00
Sarah Hoffmann	1722fc537f	bdd: add tests for non-latin scripts	2021-10-26 17:29:03 +02:00
Sarah Hoffmann	c0f347fc8c	adapt BDD tests to stricter partial search	2021-10-26 15:52:57 +02:00
Sarah Hoffmann	c4f5c11a4e	be case-insensitve about special phrase operator	2021-10-25 19:51:20 +02:00
Sarah Hoffmann	5a1c3dbea3	fix parsing of operator in special phrases Because of unstripped input, the operators wouldn't match.	2021-10-25 19:46:30 +02:00
Sarah Hoffmann	1098ab732f	allow relative paths for flatnode file	2021-10-22 17:32:51 +02:00
Sarah Hoffmann	507fdd4f40	switch IMPORT_STYLE to use generic file search Allows relative paths wrt project directory.	2021-10-22 16:49:57 +02:00
Sarah Hoffmann	0ae8d7ac08	have ADDRESS_LEVEL_CONFIG use load_sub_configuration This means that relative paths now are looked up in the project directory.	2021-10-22 16:36:52 +02:00
Sarah Hoffmann	c77df2d1eb	replace NOMINATIM_PHRASE_CONFIG with command line option	2021-10-22 14:41:14 +02:00
Sarah Hoffmann	c1fa70639b	add new replication mode catch-up This mode gets updates until the server reports no new diffs anymore. Also adds additional indexing, when the main indexing step left a couple of objects to process. This happens only when the next update is expected to be more than 40min away.	2021-10-20 22:05:15 +02:00
Sarah Hoffmann	824562357b	adapt tests for new word count mechanism	2021-10-19 12:03:48 +02:00
Sarah Hoffmann	552fb16cb2	fix template expressions for tablespaces	2021-10-15 15:11:09 +02:00
Sarah Hoffmann	3649487f5e	use SP-GIST index for building index where available Point-in-polygon queries are much faster with a SP-GIST geometry index, so use that for the index used to check if a housenumber is inside a building. Only available with Postgis 3. There is an automatic fallback to GIST for Postgis 2.	2021-10-10 21:55:38 +02:00
Sarah Hoffmann	299934fd2a	reorganize and complete tests around generic token analysis	2021-10-06 17:03:37 +02:00
Sarah Hoffmann	b18d042832	add tests for sanitizer tagging language	2021-10-06 12:29:25 +02:00
Sarah Hoffmann	97a10ec218	apply variants by languages Adds a tagger for names by language so that the analyzer of that language is used. Thus variants are now only applied to names in the specific language and only tag name tags, no longer to reference-like tags.	2021-10-06 11:09:54 +02:00
Sarah Hoffmann	d35400a7d7	use analyser provided in the 'analyzer' property Implements per-name choice of analyzer. If a non-default analyzer is choosen, then the 'word' identifier is extended with the name of the ana;yzer, so that we still have unique items.	2021-10-05 14:10:32 +02:00
Sarah Hoffmann	9ba2019470	precompute replacements while loading configuration	2021-10-05 10:20:08 +02:00
Sarah Hoffmann	7cfcbacfc7	make token analyzers configurable modules Adds a mandatory section 'analyzer' to the token-analysis entries which define, which analyser to use. Currently there is exactly one, generic, which implements the former ICUNameProcessor.	2021-10-04 17:37:34 +02:00
Sarah Hoffmann	52847b61a3	extend ICU config to accomodate multiple analysers Adds parsing of multiple variant lists from the configuration. Every entry except one must have a unique 'id' paramter to distinguish the entries. The entry without id is considered the default. Currently only the list without an id is used for analysis.	2021-10-04 16:40:28 +02:00
Sarah Hoffmann	6b348d43c6	replace test variable for PG env tests 'tty' was removed in PG14 and causes an error.	2021-10-01 12:27:24 +02:00
Sarah Hoffmann	732cd27d2e	add unit tests for new sanatizer functions	2021-10-01 12:27:24 +02:00
Sarah Hoffmann	8171fe4571	introduce sanitizer step before token analysis Sanatizer functions allow to transform name and address tags before they are handed to the tokenizer. Theses transformations are visible only for the tokenizer and thus only have an influence on the search terms and address match terms for a place. Currently two sanitizers are implemented which are responsible for splitting names with multiple values and removing bracket additions. Both was previously hard-coded in the tokenizer.	2021-10-01 12:27:24 +02:00
Sarah Hoffmann	16daa57e47	unify ICUNameProcessorRules and ICURuleLoader There is no need for the additional layer of indirection that the ICUNameProcessorRules class adds. The ICURuleLoader can fill the database properties directly.	2021-10-01 12:27:24 +02:00
Sarah Hoffmann	be65c8303f	export more data for the tokenizer name preparation Adds class, type, country and rank to the exported information and removes the rather odd hack for countries. Whether a place represents a country boundary can now be computed by the tokenizer.	2021-09-29 11:54:14 +02:00
Sarah Hoffmann	231250f2eb	add wrapper class for place data passed to tokenizer This is mostly for convenience and documentation purposes.	2021-09-29 11:54:07 +02:00
Sarah Hoffmann	40f9d52ad8	Merge pull request #2454 from lonvia/sort-out-token-assignment-in-sql ICU tokenizer: switch match method to using partial terms	2021-09-28 09:45:15 +02:00
Sarah Hoffmann	09c9fad6c3	adapt tests to new ICU address token handling	2021-09-27 17:36:23 +02:00
Sarah Hoffmann	bd7c7ddad0	icu tokenizer: switch to matching against partial names When matching address parts from addr:* tags against place names, the address names where so far converted to full names and compared those to the place names. This can become problematic with the new ICU tokenizer once we introduce creation of different variants depending on the place name context. It wouldn't be clear which variant to produce to get a match, so we would have to create all of them. To work around this issue, switch to using the partial terms for matching. This introduces a larger fuzziness between matches but that shouldn't be a problem because matching is always geographically restricted. The search terms created for address parts have a different problem: they are already created before we even know if they are going to be used. This can lead to spurious entries in the word table, which slows down searching. This problem can also be circumvented by using only partial terms for the search terms. In terms of searching that means that the address terms would not get the full-word boost, but given that the case where an address part does not exist as an OSM object should be the exception, this is likely acceptable.	2021-09-27 11:36:19 +02:00
Sarah Hoffmann	6d7c067461	force update on rank30 children when place name changes Name changes may have an effect on parenting. Don't update surrounding rank30 objects with addr:place tags as this is potentially too expensive.	2021-09-27 11:04:17 +02:00
Sarah Hoffmann	316205e455	force update of surrounding houses when street name changes When the street changes its name then this may cause changes in the parenting of rank-30 objects with an addr:street tag. Fixes #2242.	2021-09-27 10:22:41 +02:00
Sarah Hoffmann	56124546a6	fix dynamic assignment of address parts A boolean check for dynamic changes of address parts is not sufficient. The order of choice should be: 1. an addr:* part matches the name 2. the address part surrounds the object 3. the address part was declared as isaddress The implementation uses a slightly different ordering to avoid geometry checks unless strictly necessary (isaddress is false and no matching address). See #2446.	2021-09-19 12:34:39 +02:00
Sarah Hoffmann	8e1d4818ac	use yaml config loader for country info	2021-09-04 00:22:55 +02:00
Sarah Hoffmann	28c98584c1	add tests for generic YAML config reader	2021-09-03 22:31:30 +02:00
Sarah Hoffmann	1c42780bb5	introduce generic YAML config loader Adds a function to the Configuration class to load a YAML file. This means that searching for the file is generalised and works the same now for all configuration files. Changes the search logic, so that it is always possible to have a custom version of the configuration file in the project directory. Move ICU tokenizer to use new load function.	2021-09-03 18:20:07 +02:00
Sarah Hoffmann	79da96b369	read partition and languages from config file	2021-09-02 14:41:11 +02:00
Sarah Hoffmann	78fcabade8	move country name generation to country_info module	2021-09-02 14:41:11 +02:00
Sarah Hoffmann	284645f505	move generation of country tables in own module	2021-09-02 14:41:11 +02:00
Sarah Hoffmann	28ee3d0949	move linking of places to the preparation stage Linked places may bring in extra names. These names need to be processed by the tokenizer. That means that the linking needs to be done before the data is handed to the tokenizer. Move finding the linked place into the preparation stage and update the name fields. Everything else is still done in the indexing stage.	2021-08-20 22:44:17 +02:00
Sarah Hoffmann	118858a55e	rename legacy_icu tokenizer to icu tokenizer The new icu tokenizer is now no longer compatible with the old legacy tokenizer in terms of data structures. Therefore there is also no longer a need to refer to the legacy tokenizer in the name.	2021-08-17 23:11:47 +02:00
Sarah Hoffmann	5f2b9e317a	add tests for US state hacks IL, AS and LA are replaced with the US state in Geocode because the old tokenizer would simply remove the abbreviations otherwise.	2021-08-17 10:49:07 +02:00
Sarah Hoffmann	1147b83b22	php: make word list a first-class object This separates the logic of creating word sets from the Phrase class. A tokenizer may now derived the word sets any way they like. The SimpleWordList class provides a standard implementation for splitting phrases on spaces.	2021-08-16 11:51:49 +02:00
Sarah Hoffmann	87dedde5d6	allow multiple files for the import command The files are forwarded to osm2pgsql which is now able to merge them correctly.	2021-08-14 21:42:21 +02:00
Sarah Hoffmann	1db098c05d	reinstate word column in icu word table Postgresql is very bad at creating statistics for jsonb columns. The result is that the query planer tends to use JIT for queries with a where over 'info' even when there is an index.	2021-07-28 11:31:47 +02:00
Sarah Hoffmann	324b1b5575	bdd tests: do not query word table directly The BDD tests cannot make assumptions about the structure of the word table anymore because it depends on the tokenizer. Use more abstract descriptions instead that ask for specific kinds of tokens.	2021-07-28 11:31:47 +02:00
Sarah Hoffmann	e42878eeda	adapt unit test for new word table Requires a second wrapper class for the word table with the new layout. This class is interface-compatible, so that later when the ICU tokenizer becomes the default, all tests that depend on behaviour of the default tokenizer can be switched to the other wrapper.	2021-07-28 11:31:47 +02:00
Sarah Hoffmann	eb6814d74e	convert word info column to json before copying	2021-07-28 11:31:47 +02:00
Sarah Hoffmann	0c023fb4d2	adapt cli tests to Python port for add-data	2021-07-26 10:41:37 +02:00
Sarah Hoffmann	878835e4bd	move add-data subcommand into a separate file	2021-07-25 18:14:12 +02:00
Sarah Hoffmann	62d5984b1b	limit the number of variants that can be produced	2021-07-04 10:28:28 +02:00
Sarah Hoffmann	e85f7e7aa9	fix subsequent replacements Two replacement words directly following each other did not work as expected because each expects a space at the beginning/end while there was only one space available. Also forbit composing a word after a space was added in the end by a previous replacement.	2021-07-04 10:28:28 +02:00
Sarah Hoffmann	b9fbfeff67	only consider partials in multi-words for initial count This ensures that it is less likely that we exclude meaningful words like 'hauptstrasse' just because they are frequent.	2021-07-04 10:28:20 +02:00
Sarah Hoffmann	62828fc5c1	switch to a more flexible variant description format The new format combines compound splitting and abbreviation. It also allows to restrict rules to additional conditions (like language or region). This latter ability is not used yet.	2021-07-04 10:28:20 +02:00
Sarah Hoffmann	a6aa6360e0	use yaml tag syntax to mark include files	2021-07-04 10:28:20 +02:00
Sarah Hoffmann	0d80a9b897	tests for composing decomposed suffixes	2021-07-04 10:28:20 +02:00
Sarah Hoffmann	f70930b1a0	make compund decomposition pure import feature Compound decomposition now creates a full name variant on import just like abbreviations. This simplifies query time normalization and opens a path for changing abbreviation and compund decomposition lists for an existing database.	2021-07-04 10:28:20 +02:00
Sarah Hoffmann	9ff4f66f55	complete tests for icu tokenizer	2021-07-04 10:28:20 +02:00
Sarah Hoffmann	2e81084f35	complete tests for rule loader	2021-07-04 10:28:20 +02:00
Sarah Hoffmann	a0a7b05c9f	correctly quote strings when copying in data Encapsulate the copy string in a class that ensures that copy lines are written with correct quoting.	2021-07-04 10:28:20 +02:00
Sarah Hoffmann	2f6e4edcdb	update unit tests for adapted abbreviation code	2021-07-04 10:28:20 +02:00
Sarah Hoffmann	2e3c5d4c5b	adapt tests for ICU tokenizer	2021-07-04 10:28:20 +02:00
Sarah Hoffmann	8413075249	move abbreviation computation into import phase This adds precomputation of abbreviated terms for names and removes abbreviation of terms in the query. Basic import works but still needs some thorough testing as well as speed improvements during import. New dependency for python library datrie.	2021-07-04 10:28:20 +02:00
Sarah Hoffmann	e7b4fc70e7	make sure old data gets deleted on place type change When changing from some other place type to place=postcode make sure that the old place type entry in the place table is deleted.	2021-06-18 10:58:41 +02:00
Sarah Hoffmann	457982e1d2	update postcode in place if it already exists	2021-06-18 00:28:52 +02:00
Sarah Hoffmann	aa558e6080	Merge pull request #2369 from lonvia/exclude-poi-from-housenumber-search Do not return POIs when dropping house number in query	2021-06-17 15:30:05 +02:00
Sarah Hoffmann	fe11d3cbbd	do not return POIs when dropping house number in query We've previously added searching through rank 30 in a house number search to enable searches for house number+name. This had the unintended side effect that rank 30 objects are also returned in s search that dropped the house number from the query. This is wrong because POIs cannot function as a parent to a house number. This fix drops all rank 30 objects from the results for a house number search if they do not match the requested house number.	2021-06-17 14:21:20 +02:00
AntoJvlt	3676310efe	Improved performance of the postcodes query and some code cleaning	2021-06-12 15:46:08 +02:00
AntoJvlt	1c175e3a67	Clean and update tests for postcodes	2021-06-09 09:31:32 +02:00
AntoJvlt	e879814e43	Update tests for postcodes	2021-06-09 09:31:32 +02:00
Sarah Hoffmann	3aac51c81f	switch BDD tests to always use search API	2021-06-06 15:27:52 +02:00
Sarah Hoffmann	bc981d0261	fix insertion of special terms and countries into word table Special terms need to be prefixed by a space because they are full terms. For countries avoid duplicate entries of word tokens. Adds tests for adding country terms.	2021-06-02 20:22:39 +02:00
Sarah Hoffmann	24c986c842	add tests for new full name computation with ICU	2021-05-24 10:41:42 +02:00
Sarah Hoffmann	4f4d15c28a	reorganize keyword creation for legacy tokenizer - only save partial words without internal spaces - consider comma and semicolon a separator of full words - consider parts before an opening bracket a full word (but not the part after the bracket) Fixes #244.	2021-05-24 10:41:42 +02:00
Sarah Hoffmann	10143e0ac7	Merge pull request #2342 from lonvia/icu-tokenizer-ci Add BDD tests with icu tokenizer to CI runs	2021-05-22 10:36:35 +02:00
Sarah Hoffmann	00094c43d1	enable Tiger BDD API test for legacy_icu	2021-05-21 22:39:56 +02:00
Sarah Hoffmann	430c316e45	test: fix linting errors	2021-05-19 23:07:39 +02:00
Sarah Hoffmann	01f5a9ff84	test: more use of table_factory	2021-05-19 17:37:03 +02:00
Sarah Hoffmann	af52eed0dd	test: avoid use of tempfile module Use the tmp_path fixture instead which provides automatic cleanup.	2021-05-19 16:43:26 +02:00
Sarah Hoffmann	f93d0fa957	test: use src_dir fixture instead of self-computed paths	2021-05-19 16:03:54 +02:00
Sarah Hoffmann	c06a1d007a	test: replace raw execute() with fixture code where possible	2021-05-19 12:11:04 +02:00
Sarah Hoffmann	65bd749918	test: use table_rows() and execute_values() where possible Some uses of scalar() could also be replaced with convenience functions from the word table mock.	2021-05-19 10:51:10 +02:00
Sarah Hoffmann	510eb53f53	test: move Testingcursor into separate class Also adds more convenience functions: counting with a where statement and a wrapper to execute_values().	2021-05-19 10:30:36 +02:00
Sarah Hoffmann	16bb007135	Merge pull request #2336 from lonvia/do-not-mask-error-when-loading-tokenizer Do not hide errors when importing tokenizer	2021-05-18 23:00:10 +02:00
Sarah Hoffmann	b2722650d4	do not hide errors when importing tokenizer Explicitly check for the tokenizer source file to check that the name is correct. We can't use the import error for that because it hides other import errors like a missing library. Fixes #2327.	2021-05-18 16:28:21 +02:00
AntoJvlt	3206bf59df	Resolve conflicts	2021-05-17 13:52:35 +02:00
AntoJvlt	8b8dfc46eb	Added --no-replace command for special phrases importation and added corresponding tests	2021-05-17 13:25:06 +02:00
AntoJvlt	06aab389ed	Code cleaning and SPLoader deleted	2021-05-16 16:59:12 +02:00
AntoJvlt	fb0ebb5bf0	Add tests for the new SPWikiLoader and SPCsvLoader	2021-05-16 16:10:06 +02:00
Sarah Hoffmann	925726222f	Merge pull request #2323 from darkshredder/disable-search-reverse-only Feat: Disabled search API for --reverse-only imports	2021-05-14 10:40:22 +02:00
Sarah Hoffmann	7d621389ee	adapt tests to new TIGER CSV format	2021-05-14 00:02:50 +02:00
Darkshredder	e5ffc59cd5	feat: Added reverse-only-search validation	2021-05-14 02:36:21 +05:30
Sarah Hoffmann	5feece64c1	use WorkerPool for Tiger data import Requires adding an option that SQL errors are ignored.	2021-05-13 20:36:50 +02:00
Sarah Hoffmann	f5977dac75	ignore invalid coordinates in external postcodes	2021-05-13 14:15:42 +02:00
Sarah Hoffmann	8f2746fe24	ignore entries without country code	2021-05-13 14:15:42 +02:00
Sarah Hoffmann	1ccd4360b4	correctly handle removing all postcodes for country	2021-05-13 14:15:42 +02:00
Sarah Hoffmann	bf864b2c54	index postcodes after refreshing	2021-05-13 14:15:42 +02:00
Sarah Hoffmann	4abaf71234	add and extend tests for new postcode handling	2021-05-13 14:15:42 +02:00
AntoJvlt	9d83da830f	Introduction of SPCsvLoader to load special phrases from a csv file	2021-05-10 23:26:39 +02:00
AntoJvlt	00959fac57	Refactoring loading of external special phrases and importation process by introducing SPLoader and SPWikiLoader	2021-05-10 21:49:31 +02:00
Sarah Hoffmann	b2c6eca2c8	add missing transliterations The ICU library only offers transliterations for a limited set of script. Add transliterations for missing scripts from the PostgreSQL module. These means that the same selection of scripts is supported as with the old module.	2021-05-05 21:16:55 +02:00
Sarah Hoffmann	a263e54b94	enable BDD tests for different tokenizers The tokenizer to be used can be choosen with -DTOKENIZER. Adapt all tests, so that they work with legacy_icu tokenizer. Move lookup in word table to a function in the tokenizer. Special phrases are temporarily imported from the wiki until we have an implementation that can import from file. TIGER tests do not work yet.	2021-05-05 10:31:51 +02:00
Sarah Hoffmann	18c99a5c5f	add unit tests for legacy ICU tokenizer	2021-05-05 10:15:27 +02:00
Sarah Hoffmann	8bdb9aa607	mock tokenizer factory for replication tests	2021-05-01 10:50:39 +02:00
Sarah Hoffmann	388ebcbae2	move index creation for word table to tokenizer This introduces a finalization routing for the tokenizer where it can post-process the import if necessary.	2021-04-30 17:41:08 +02:00
Sarah Hoffmann	fc995ea6b9	move database check for module to tokenizer	2021-04-30 17:41:08 +02:00
Sarah Hoffmann	be6262c6ce	move status test to tokenizer The availability of the module is now tested by the tokenizer.	2021-04-30 17:41:08 +02:00
Sarah Hoffmann	893490f94e	add more tests for legacy tokenizer	2021-04-30 17:41:08 +02:00

... 3 4 5 6 7 ...

877 Commits