Software
I develop and maintain a number of free software projects, most of them related to my research or to scientific publishing. Most are hosted on GitHub.
Research software
- MIT ProvSQL: support for (semiring) provenance and uncertainty management in PostgreSQL databases [publication]
- MIT provenance-lean: Lean 4 formalization of some notions of data provenance [publication]
- Apache descriptive-complexity: Lean 4 formalization of descriptive complexity theory, enabling machine-model-free hardness reductions in the style of Immerman
- MIT TheoremKB: collection of tools to extract semantic information from (mathematical) research articles [publication]
- MIT lsg: large sparse graph library [publication]
LaTeX and scientific publishing
- LPPL proofgraph: LaTeX package automatically producing a graph of the dependencies between the results of a mathematical article
- Apache result-graph: counterpart of proofgraph for Lean 4, producing graphs of the dependencies between the results of a formalization, sliced per declaration
- LPPL apxproof: LaTeX package for automatically typesetting proofs (and other material) in appendix
- MIT dblpify: using DBLP to clean up and make uniform the references in a BibTeX file
Miscellaneous utilities
- MIT muttlike-imap: searching IMAP mailboxes from the command line using mutt-style patterns
- MIT recletters: Web system for uploading recommendation letters
Archived software
The following software is no longer maintained. It is kept online for the record, as the software behind past publications.
- LPPL erc-latex-template: LaTeX template for ERC grant proposals
- focusedrl: focused crawling with reinforcement learning, developed by Miyoung Han during her PhD [publication]
- taxiroute: reinforcement-learning taxi routing, developed by Miyoung Han during her PhD [publication]
- AGPL Dissemin: Web platform helping researchers deposit their publications in open repositories (service now discontinued)
- FOREST: extraction of the main content of Web pages, and the RED dataset of annotated Web pages, developed by Marilena Oita during her PhD [publication]
- GPL AAH: application-aware helper for the crawling of Web applications (forums, blogs, content management systems) for Web archiving, and ACE adaptive crawler, developed by Muhammad Faheem during his PhD [publication]
- ProApproX: approximation query processor over probabilistic XML documents, developed by Asma Souihli during her PhD [publication]
- datacorrob: implementation and datasets for truth finding (data corroboration) [publication]
- hiddenweb: probing of hidden-Web forms and induction of wrappers for their result pages, with the material of the experiments (2008) [publication]
- MIT Wikipedia scripts: scripts to extract content (in particular, the graph structure) from Wikipedia [publication]
- MIT Fuzzy XML: Java implementation of a probabilistic XML model [publication]
- websiteid: crawler and identification of logical Web sites by flow simulation, in Java (2003) [publication]
- mulverif: formal verification of multiplier circuits with binary moment diagrams, in Haskell (2002) [publication]