Utils#
Various functions and classes.
Common#
All-purpose functions.
- class gismap.utils.common.Data(data)[source]#
Easy-going converter of dict to dataclass. Useful when you want to use attribute access and do not care about giving a full description.
Examples
>>> data = Data({ ... 'name': 'Alice', ... 'age': 30, ... 'address': {'street': '123 Main', 'city': 'Paris'}, ... 'hobbies': [{'name': 'jazz', 'level': 5}, {'name': 'code'}]}) >>> data Data(name='Alice', age=30, address=Data(street='123 Main', city='Paris'), hobbies=[Data(name='jazz', level=5), Data(name='code')]) >>> data.hobbies[0].name 'jazz' >>> data.todict() {'name': 'Alice', 'age': 30, 'address': {'street': '123 Main', 'city': 'Paris'}, 'hobbies': [{'name': 'jazz', 'level': 5}, {'name': 'code'}]}
- class gismap.utils.common.LazyRepr[source]#
MixIn that provides a clean repr for dataclasses.
Hides empty fields and fields in HIDDEN_KEYS from the repr string. Private attributes (starting with ‘_’) are also hidden.
- gismap.utils.common.list_of_objects(clss, dico, default=None)[source]#
Versatile way to enter a list of objects referenced by a dico.
- Parameters:
- Returns:
Proper list of objects.
- Return type:
Examples
>>> from gismap.sources.models import DB >>> from gismap import HAL, DBLP, LDB # force registration >>> subclasses = get_classes(DB, key='db_name') >>> list_of_objects([HAL, 'ldb'], subclasses) [<class 'gismap.sources.hal.HAL'>, <class 'gismap.sources.ldb.LDB'>] >>> list_of_objects(None, subclasses, [DBLP]) [<class 'gismap.sources.dblp.DBLP'>] >>> list_of_objects(LDB, subclasses) [<class 'gismap.sources.ldb.LDB'>] >>> list_of_objects('hal', subclasses) [<class 'gismap.sources.hal.HAL'>]
Requests#
Functions related to the requests.
- gismap.utils.requests.get(url, params=None, n_trials=10, verify=True, encoding=None, timeout=(10, 30))[source]#
- Parameters:
url (
str) – Entry point to fetch.params (
dict, optional) – Get arguments (appended to URL).n_trials (
int, default=10) – Number of attempts to fetch URL.verify (
bool, default=True) – Verify certificates.encoding (
str, optional) – Force response encoding (e.g."utf-8"). Useful when the server does not declare the charset andrequestsfalls back to ISO-8859-1.timeout (
floatortuple, default=(10, 30)) –(connect, read)timeout in seconds passed torequests. Without it a slow or hung server blocks forever; a timeout turns that into a retry.
- Returns:
Result.
- Return type:
Logger#
Keep track of things.
- gismap.utils.logger.logger = <Logger GisMap (INFO)>#
Default logging interface.
Zlist#
Convert a list into a succession of compressed frames. Reduces memory footprint at the price of slower random access (sequential access is unaffected).
- class gismap.utils.zlist.ZList(frame_size=1000, level=3, dict_data=None)[source]#
Compressed list with frame-based storage.
Stores elements in compressed frames, allowing efficient memory usage while maintaining random access. Uses zstandard compression.
In this version, each frame is a concatenation of individually pickled items prefixed by an intra-frame offset index, then compressed. Random access therefore unpickles a single item instead of the whole frame, and an optional zstd dictionary can be trained for tighter compression.
Typical use is a two-pass build: stream items through the default constructor (fast, no dictionary), then call
optimize()to train a dictionary and recompress aggressively (or fall back to a plain list when the data is small enough that memory is not a concern).Use as a context manager for building:
- with ZList(frame_size=100) as z:
- for item in data:
z.append(item)
Or use the
from_iterable()method.- Parameters:
Examples
Let us build a small big list:
>>> mylist = [c * 1000 for c in "abcdefghijklmnopqrstuvwxyz"]
One builds a ZList out of it.
>>> zlist = ZList.from_iterable(mylist, frame_size=10)
Why ZLists? Because sometimes size matters: the compressed blob is far smaller than the raw data.
>>> raw_bytes = sum(len(s) for s in mylist) >>> raw_bytes 26000 >>> 0 < len(zlist._blob) < raw_bytes True
>>> zlist[20][-10:] 'uuuuuuuuuu' >>> len(zlist) 26 >>> for i, line in enumerate(mylist): ... assert zlist[i] == line
A ZList can also be obtained using a context manager and successive append.
>>> with ZList(frame_size=10, level=0) as zlist2: ... for line in mylist: ... zlist2.append(line)
Once built, the list can be packed into its best storage form with
optimize(). A small source is returned as a plainlist(memory is cheap, and decompressed access is faster):>>> isinstance(zlist.optimize(), list) True
A tight memory budget keeps it compressed as a ZList instead:
>>> isinstance(zlist.optimize(max_bytes=1000), ZList) True
- estimated_uncompressed_size(fudge=4)[source]#
Estimate the in-RAM footprint (in bytes) of the decompressed items.
The uncompressed size of each zstd frame is read straight from its header (no decompression required), then scaled by fudge to account for the overhead of live Python objects compared to their pickled bytes. For the dataclass items that flow through a source, the live footprint measures about 3.8x the pickled bytes, so the default of 4 keeps the estimate a safe upper bound; deeply nested payloads may need a larger value.
- optimize(frame_size=10, level=19, threshold=10000, max_bytes=10000000)[source]#
Return the source in its best storage form for its size.
When the estimated decompressed footprint is below max_bytes, the items are returned as a plain
list: memory is not a concern at that size, and a decompressed list avoids the per-access decompression and unpickling cost of a ZList. Otherwise the list is rebuilt as a compressed ZList using the provided frame_size / level. When dict_data is not set and the list contains more than threshold items, a zstd dictionary is trained to improve compression.- Parameters:
frame_size (
int, default=10) – Number of elements per compressed frame (ZList path only).level (
int, default=19) – Level of compression (ZList path only).threshold (
int, default=10000) – Train a (missing) dictionary only above this size threshold (in items).max_bytes (
int, default=10_000_000) – Below this estimated decompressed footprint, return a plain list.
- Returns:
A plain list for small sources, a recompressed ZList otherwise.
- Return type:
Text#
Text manipulation tools.
- class gismap.utils.text.Corrector(voc, score_cutoff=20, min_length=3)[source]#
A simple word corrector base on input vocabulary. Short words are discarded.
- Parameters:
Examples
>>> vocabulary = ['My Taylor Swift is Rich'] >>> phrase = "How riche ise Tailor Swyft" >>> cor = Corrector(vocabulary, min_length=4) >>> cor(phrase) 'How rich ise taylor swift' >>> cor = Corrector(vocabulary, min_length=2) >>> cor(phrase) 'How rich is taylor swift'
- gismap.utils.text.asciify(text)[source]#
- Parameters:
text (
str) – Some text (typically names) with annoying accents.- Returns:
Same text simplified into ascii.
- Return type:
Examples
>>> asciify('Ana Bušić') 'Ana Busic' >>> asciify("Thomas Deiß") 'Thomas Deiss'
- gismap.utils.text.normalized_name(txt)[source]#
Try to normalize names for facilitating comparisons. Name is lowered, split, asciified, sorted, and filtered.
Examples
>>> normalized_name("Thomas Deiß") 'deiss thomas' >>> normalized_name("Dario Rossi 001") 'dario rossi' >>> normalized_name("James W. Roberts") 'james roberts'
- gismap.utils.text.normalized_title(txt)[source]#
Try to normalize titles for facilitating comparisons. Title is lowercased, asciified, and stripped of punctuation.
Examples
>>> normalized_title("An Efficient Algorithm for P2P Networks") 'an efficient algorithm for p2p networks' >>> normalized_title("A Study on the Use of Millimeter Waves: 5G Networks.") 'a study on the use of millimeter waves 5g networks'
Fuzzy#
Fuzzy matching utilities (similarity_matrix).
- gismap.utils.fuzzy.similarity_matrix(references, candidates=None, n_range=4, length_impact=0.05, key=None, key2=None)[source]#
Compute a similarity matrix between objects using fuzzy n-gram matching.
When
candidatesis None, computes pairwise similarities withinreferences(self-comparison). Whencandidatesis provided, computes cross-similarities betweenreferencesandcandidates.- Parameters:
references (
list) – Reference objects.candidates (
list, optional) – Candidate objects to compare against references. If None, references are compared against themselves.n_range (
int, default=4) – N-gram range for the vectorizer.length_impact (
float, default=0.05) – Impact of length difference on similarity scores.key (callable, optional) – Fingerprint extractor for references. Defaults to identity.
key2 (callable, optional) – Fingerprint extractor for candidates. Defaults to
key.
- Returns:
Similarity matrix. Shape is
(len(references), len(references))for self-comparison, or(len(candidates), len(references))for cross-comparison.- Return type:
Examples
>>> m = similarity_matrix(["abc def", "abc deg", "xyz"]) >>> m.shape (3, 3) >>> m[0, 1] > 50 True >>> m[0, 2] < 50 True