Laboratory#
Management of a group of people and their publications is made with the LabMap abstract class.
LabMaps#
- class gismap.lab.labmap.LabMap(name=None, dbs=None, *, max_co_authors=9, min_title_words=2, taboo_words=None, taboo_authors=None)[source]#
Abstract class for labs.
Actual Lab classes can be created by implementing the _author_iterator method.
Labs can be saved with the dump method and loaded with the load method.
- Parameters:
name (
str) – Name of the lab. Can be set as class or instance attribute.max_co_authors (
int, default=9) – Reject publications with strictly more than this number of co-authors.min_title_words (
int, default=2) – Reject publications whose title contains fewer than this number of words.taboo_words (
list, optional) – Words/regexes; a publication is rejected if its title matches any of them. Defaults toeditorials.taboo_authors (
list, optional) – Words/regexes; an author is rejected if their name matches any of them. Defaults tocharlatans.
- author_selectors#
Author filters. Built from
taboo_authorsat construction time but can be reassigned for advanced use cases.- Type:
- publication_selectors#
Publication filters. Built from
max_co_authors,taboo_wordsandmin_title_wordsat construction time but can be reassigned.- Type:
- add_publication(title, authors, **kwargs)[source]#
Add a manual publication to the lab.
Author names given as strings are resolved to known authors from the lab’s publications using fuzzy matching. Unmatched names become
Outsiderinstances.- Parameters:
title (
str) – Publication title.**kwargs – Passed to
Informal(venue,type,year,key,metadata) and tofit_names()(threshold,n_range,length_impact).
- del_publication(query, confirm=True, **kwargs)[source]#
Remove publications matching a query from the lab.
- Parameters:
query (
strorcallable) – Passed toselect_publications().confirm (
bool, default=True) – If True, display matches and prompt for confirmation before deletion.**kwargs – Passed to
select_publications().
- expand(target=None, group='moon', desc='Moon information', **kwargs)[source]#
Expand the lab with external collaborators found in publications.
Discovers authors who co-published with lab members, ranks them by collaboration strength, and adds the top candidates.
- Parameters:
target (
int, optional) – Number of new authors to add. Defaults tolen(self.authors) // 3.group (
str, default=”moon”) – Group label assigned to new authors.desc (
str, default=”Moon information”) – Progress bar description.**kwargs – Passed to
proper_prospects().
- gismo_lab(**kwargs)[source]#
Build a
GismoLab(authors ↔ keywords cross-embedding) from this lab. Reuse a single instance to run several keyword/wordcloud queries without rebuilding the embeddings.- Parameters:
**kwargs – Passed to
GismoLab(ngram_range,stop_words).
- html(**kwargs)[source]#
Generate HTML representation of the collaboration graph.
- Parameters:
**kwargs – Passed to
make_vis().- Returns:
HTML content as a string.
- Return type:
- keywords(query=None, *, gismo_lab=None, **kwargs)[source]#
Ranked
(word, weight)keywords for a topic, an author, a group, or the whole lab. Builds a freshgismo_lab()each call unless one is supplied; reuse a singlegismo_lab()for many queries to avoid rebuilding.- Parameters:
query (
strorlist, optional) – A text query, or a list of author keys. Defaults to the whole lab.gismo_lab (
GismoLab, optional) – A prebuilt instance to reuse; ifNone, a fresh one is built.**kwargs – Ranking tuning forwarded to
keywords()(e.g.groupto restrict to a group,k,threshold).
- propose_overrides(**kwargs)[source]#
Propose
overridesfor members whose name matches several LDB entries, ranking the candidates by the co-authors they share with the rest of the lab. Nothing is applied: print the result, check it, then copy what you agree with intooverridesand runupdate_authors()again.- Parameters:
**kwargs – Passed to
propose_overrides().- Return type:
- regroup(pubs=None)[source]#
Rebuild the publication clusters from their raw sources.
Existing clusters are flattened back to their sources, merged with the optional extra raw publications, and clustered again after re-attaching the lab authors. Call it once after a batch of
add_publication(): a manual entry is stored as a cluster of its own, so it only merges with its database counterparts (e.g. the preprint of an accepted paper) on the next regroup. Regrouping after each addition would be quadratic, hence the explicit call.- Parameters:
pubs (
dict, optional) – Extra raw publications (source key -> publication) to merge before clustering.- Return type:
None
- save_html(name=None, **kwargs)[source]#
Save the collaboration graph as an HTML file.
- Parameters:
name (
str, optional) – Output filename. Defaults to lab name. RaisesValueErrorif neither is set.**kwargs – Passed to
html().
- Return type:
None
- select_publications(query, n_range=4, length_impact=0.001, threshold=80)[source]#
Search for publications matching a query.
- Parameters:
query (
strorcallable) – If a string, matches by exact key or fuzzy title similarity. If a callable, used as a filterf(pub) -> boolon each publication.n_range (
int, default=4) – Passed tosimilarity_matrix().length_impact (
float, default=0.001) – Passed tosimilarity_matrix().threshold (
int, default=80) – Minimum similarity score (0-100) for fuzzy title matching.
- Returns:
Matching publications.
- Return type:
- show_html(**kwargs)[source]#
Display the collaboration graph in a Jupyter notebook.
- Parameters:
**kwargs – Passed to
html().- Return type:
None
- to_bib(name=None, query=None, **kwargs)[source]#
Export publications as a BibTeX file.
- Parameters:
name (
str, optional) – Output filename. Defaults to lab name;.bibsuffix is added.query (
strorcallable, optional) – Filter for publications. Passed toselect_publications(); if None, all publications are exported.**kwargs – Forwarded to
select_publications()whenqueryis given.
- Return type:
None
- to_csv(name=None)[source]#
Export the lab to two CSV files.
Writes
<name>_authors.csv(key, name, group, url, sources) and<name>_publications.csv(cite_key, title, year, type, venue, authors, primary_url, abstract, other_urls). Multi-valued cells use|as a separator.- Parameters:
name (
str, optional) – Base filename. Defaults to lab name.- Return type:
None
- to_json(name=None)[source]#
Export the lab as a JSON file.
The structure exposes
name,authors(list ofto_dict()) andpublications(list ofto_dict()).- Parameters:
name (
str, optional) – Output filename. Defaults to lab name;.jsonsuffix is added.- Return type:
None
- update_authors(desc='Author information')[source]#
Populate the authors attribute (
dict[str,LabAuthor]).- Return type:
None
- update_publis(desc='Publications information')[source]#
Populate the publications attribute (
dict[str,SourcedPublication]).- Return type:
None
- wordcloud(query=None, *, gismo_lab=None, **kwargs)[source]#
A renderable
WordCloud(displays inline in notebooks). Same arguments askeywords().
EgoMaps#
- class gismap.lab.egomap.EgoMap(star, *args, **kwargs)[source]#
Egocentric view of a researcher’s collaboration network.
Displays the star (central researcher), their planets (direct co-authors), and optionally moons (co-authors of co-authors).
- Parameters:
Examples
>>> dang = EgoMap("The-Dang Huynh") >>> dang.build(target=20) >>> sorted( ... a.name for a in dang.authors.values() if len(a.name.split()) < 3 ... ) ['Bruno Kauffmann', 'Diego Perino', 'Dohy Hong', 'Fabien Mathieu', 'François Baccelli',...]
To add publications, one can use the
add_publication()method:>>> dang.add_publication( ... title="A new paper", ... authors=[dang.star, "Fabien Mathieu", "Alice Smith"], ... venue="Journal of Testing", ... ) >>> str(dang.select_publications(lambda p: "Testing" in p.venue)[0]) 'A new paper, by The-Dang Huynh, Fabien Mathieu, and Alice Smith. In Journal of Testing [unpublished], 2026.'
To remove publications, one can use the
del_publication()method:>>> dang.del_publication("A new paper", confirm=False) >>> dang.select_publications(lambda p: "Testing" in p.venue) []
- build(target=50, moon_ratio=0.5, **kwargs)[source]#
Build the ego network by fetching publications and adding planets/moons.
- Parameters:
target (
intor None, default=50) – Target number of authors in the final map. UseNonefor an exhaustive map: every planet (direct co-author) is kept, then moons are capped atmoon_ratiotimes the number of planets found. The cap matters because taking all moons routinely reaches several thousand authors.moon_ratio (
float, default=0.5) – Only used whentarget=None: number of moons to add, as a fraction of the number of planets found.**kwargs – Passed to
expand().
- Return type:
None
Utilities#
Expansion#
- class gismap.lab.expansion.ProspectStrength(coauthors: int, publications: int)[source]#
Measures the interaction between an external author and a lab by counting co-authors and publications.
A (max,+) addition is handled to deal with multiple keys.
Examples
>>> a1 = ProspectStrength(3, 5) >>> a2 = ProspectStrength(2, 10) >>> a1 > a2 True >>> a1 + a2 ProspectStrength(coauthors=3, publications=15)
- gismap.lab.expansion.count_prospect_entries(lab)[source]#
Associate to external coauthors (prospects) their lab strength.
- Parameters:
lab (
LabMap) – Reference lab.- Returns:
Lab strengths.
- Return type:
dictofstrtoProspectStrength
- gismap.lab.expansion.proper_prospects(lab, length_impact=0.05, threshold=80, n_range=4, max_new=None, trim=True)[source]#
Find and rank external collaborators for potential lab expansion.
Identifies authors from publications who are not already lab members, groups them by name similarity, and ranks by collaboration strength.
- Parameters:
lab (
LabMap) – Reference lab.length_impact (
float, default=0.05) – Length impact for name similarity matching.threshold (
int, default=80) – Similarity threshold for grouping authors.n_range (
int, default=4) – N-gram range for name comparison.max_new (
int, optional) – Maximum number of new authors to return.trim (
bool, default=True) – If True, keep only one source per database for each author.
- Returns:
New authors ranked by collaboration strength (descending).
- Return type:
- gismap.lab.expansion.trim_sources(author)[source]#
Inplace reduction of sources, keeping one unique source per db.
- Parameters:
author (
SourcedAuthor) – An author.- Return type:
None
Disambiguation#
Semi-automatic disambiguation of homonyms.
When a member’s name matches several entries of the same database (typically the
-N-suffixed entries of DBLP/LDB), auto_sources()
keeps them all, so the publications of unrelated homonyms end up merged into one
author. The tools below rank the candidates by the number of co-authors they share
with the rest of the lab and propose an override spec for the ones that stand out.
In the spirit of HALTools, they surface entries to check, not entries to change:
nothing is applied until the proposals are copied into
overrides.
Only LDB is supported: DBLP online is rate-limited (several seconds per author, so inspecting dozens of candidates is impractical) and HAL’s multiple keys designate the same person by construction, so there is nothing to disambiguate there.
- gismap.lab.disambiguation.AMBIGUOUS_DBS = ('dblp', 'ldb')#
Databases where several entries matching one name designate distinct persons.
- class gismap.lab.disambiguation.Candidate(key: str, shared: int, total: int, discarded: bool = False)[source]#
One LDB entry matching an ambiguous author name.
- class gismap.lab.disambiguation.OverrideProposals(proposals: list)[source]#
Result of
propose_overrides(): oneProposalper ambiguous author.Print it to review the candidates;
to_dict()gives the overrides of the decided proposals, ready forlab.overrides.update(...).
- class gismap.lab.disambiguation.Proposal(name: str, candidates: list, keys: list)[source]#
Candidates of one ambiguous author, and the keys proposed for pinning.
- Parameters:
- gismap.lab.disambiguation.ldb_coauthors(index)[source]#
Co-authors of an LDB author, computed in index space: no author or publication object is built, which keeps catch-all entries with thousands of co-authors cheap.
- gismap.lab.disambiguation.propose_overrides(lab, max_co_authors=1000)[source]#
Propose overrides for the members of a lab whose name matches several LDB entries.
For each such member, every candidate entry is scored by the number of co-authors it shares with the rest of the lab (the member’s own candidates excluded). The proposal is the set of candidates reaching the best score, so the function never picks arbitrarily: when two entries tie, both are proposed (DBLP sometimes splits one person in two, and keeping both is then right). When no candidate shares any co-author, or when all of them tie, the author is reported as undecided and nothing is proposed, which leaves the default multi-source behavior untouched.
Entries with more than
max_co_authorsco-authors are set aside before scoring: for very common names, the un-suffixed DBLP entry is not the main person but the bin where DBLP parks what it has not disambiguated yet.- Parameters:
lab (
LabMap) – Lab whose authors have been populated withupdate_authors().max_co_authors (
int, default=1000) – Candidates with more co-authors than this are discarded.
- Return type:
Examples
Typical workflow (network and LDB required, hence not executed here):
>>> lab.update_authors() >>> proposals = lab.propose_overrides() >>> print(proposals) === 2 ambiguous author(s), 1 proposal(s) === Bo Li: propose ldb:50/3402-37, no_auto (181 candidates, 2 discarded) ldb:50/3402-37: 6 shared / 130 co-authors ldb:50/3402-115: 1 shared / 106 co-authors ldb:50/3402-23: 1 shared / 74 co-authors Tao Lin: undecided (27 candidates) >>> lab.overrides.update(proposals.to_dict()) >>> lab.update_authors()
Filters#
- gismap.lab.filters.author_taboo_filter(w=None)[source]#
- Parameters:
w (
list, optional) – List of words to filter.- Returns:
Filter function on authors.
- Return type:
Callable
- gismap.lab.filters.publication_oneword_filter(n_min=2)[source]#
- Parameters:
n_min (int, default=2) – Minimum number of words required in the title.
- Returns:
Filter on number of words required in the title.
- Return type:
callable
- gismap.lab.filters.publication_size_filter(n_max=9)[source]#
- Parameters:
n_max (int, default=9) – Maximum number of co-authors allowed.
- Returns:
Filter on number of co-authors.
- Return type:
callable