Eng Fra

Pierre Senellart

  • Home
  • Resume
  • Publications
  • Talks
  • Teaching
  • Students
  • Software
  • Other

Research artifacts

For each publication that involves an implementation, a dataset, or formal proofs, this page lists what is available and what is not, and why. Papers that are purely theoretical, surveys, or position papers are not concerned.

  • 114 publications concerned
  • 80 fully available
  • 5 partly available
  • 29 with none available
  • 90 publications not concerned (theory, surveys, position papers)

Everything available (80)

  • datasetStructured Prediction for Scalable Spreadsheet Table Understanding: From Cell Types to Table RangesCIKM 2026
    codedata
  • demonstrationPreserving LLM Output Distributions for Probabilistic Query EvaluationCIKM 2026
    code
  • experimental studyStructured Prediction for Scalable Spreadsheet Table Understanding: From Cell Types to Table Ranges (Extended Version)arXiv 2026
    codedata
  • datasetBenchmarking Table Extraction from Heterogeneous Scientific PDF DocumentsKDD 2026
    codedatadata
  • datasetIn a Streaming World, Should You Stand Still? A Comprehensive Benchmark of Anomaly Detection in StreamsKDD 2026
    codecode
  • datasetBenchmarking Table Extraction from Heterogeneous Scientific PDF Documents (Extended Version)arXiv 2026
    data
  • experimental studyEfficient and Scalable Search for StatisticsICDE 2026
    data
  • systemProvSQL: A General System for Keeping Track of the Provenance and Probability of DataICDE 2026
    codedatadataformal proofs
  • systemEfficient Crawling for Scalable Web Data AcquisitionEDBT 2026
    code
  • systemEfficient Crawling for Scalable Web Data Acquisition (Extended Version)arXiv 2026
    code
  • systemSTAR: Efficient, Scalable, and Modular Retrieval of Statistical TablesTransactions on Large-Scale Data and Knowledge-Centered Systems 2026
    codedata
  • experimental studyDiscovering Voting Power for Ensemble MethodsDEXA 2025
    code
  • demonstrationDemonstration of ProvSQL Update Provenance through Temporal DatabasesPW 2025
    codedata
  • demonstrationTheoremView: A Framework for Extracting Theorem-Like Environments from Raw PDFsECIR 2025
    code
  • experimental studyApprentissage multimodal modulaire pour l’extraction de théorèmes et de preuves dans des documents scientifiques longsEGC 2025
    code
  • experimental studyMSAD: A Deep Dive into Model Selection for Time Series Anomaly DetectionVLDB Journal 2025
    code
  • experimental studyModular Multimodal Machine Learning for Extraction of Theorems and Proofs in Long Scientific DocumentsJCDL 2024
    code
  • experimental studyModular Multimodal Machine Learning for Extraction of Theorems and Proofs in Long Scientific Documents (Extended Version)arXiv 2024
    code
  • datasetFirst Steps in Building a Knowledge Base of Mathematical ResultsSDP 2024
    codecode
  • experimental studyExtracting Definienda in Mathematical Scholarly Articles with TransformersWIESP 2023
    codedata
  • experimental studyConfidential Truth Finding with Multi-Party ComputationDEXA 2023
    code
  • experimental studyAutomatically Inferring the Document Class of a Scientific ArticleDocEng 2023
    code
  • systemConfidential Truth Finding with Multi-Party Computation (Extended Version)arXiv 2023
    code
  • systemEfficient Provenance-Aware Querying of Graph Databases with DatalogBDA 2022
    code
  • experimental studyEfficient Provenance-Aware Querying of Graph Databases with DatalogGRADES-NDA 2022
    code
  • systemA Practical Dynamic Programming Approach to Datalog Provenance ComputationarXiv 2021
    code
  • experimental studyTowards Extraction of Theorems and Proofs in Scholarly ArticlesDocEng 2021
    code
  • experimental studyProvenance-Based Algorithms for Rich Queries over Graph DatabasesEDBT 2021
    code
  • experimental studyAlgorithmes à base de provenance pour des requêtes enrichies sur les bases de données graphesBDA 2020
    code
  • experimental studyBelMan: An Information-Geometric Approach to Stochastic BanditsECML/PKDD 2019
    code
  • experimental studyAn Experimental Study of the Treewidth of Real-World Graph DataICDT 2019
    codedata
  • experimental studyAn Experimental Study of the Treewidth of Real-World Graph Data (Extended Version)arXiv 2019
    codedata
  • demonstrationProvSQL : Gestion de provenance et de probabilités dans PostgreSQLBDA 2018
    code
  • experimental studyUne étude expérimentale de la largeur d'arbre de données graphe du monde réelBDA 2018
    codedata
  • demonstrationProvSQL: Provenance and Probability Management in PostgreSQLVLDB 2018
    code
  • experimental studySemiring Provenance over Graph DatabasesTaPP 2018
    code
  • systemForm Filling based on Constraint SolvingICWE 2018
    codedata
  • systemA Knowledge Base for Personal Information ManagementLDOW 2018
    codecodecode
  • systemProvenance and Probabilities in Relational DatabasesSIGMOD Record 2017
    code
  • systemThymeflow, An Open-Source Personal Knowledge Base SystemMiscellaneous 2016
    codecodecode
  • experimental studyRegularized Cost-Model Oblivious Database Tuning with Reinforcement LearningTransactions on Large-Scale Data and Knowledge-Centered Systems 2016
    code
  • systemAdaptive Web Crawling through Structure-Based Link ClassificationICADL 2015
    code
  • experimental studyApprentissage par renforcement pour optimiser les bases de donnéees indépendamment du modèle de coûtBDA 2015
    code
  • experimental studyCost-Model Oblivious Database Tuning with Reinforcement LearningDEXA 2015
    code
  • experimental studyOnline Influence MaximizationKDD 2015
    code
  • experimental studyOnline Influence Maximization (Extended Version)arXiv 2015
    code
  • systemFOREST: Focused Object Retrieval by Exploiting Significant Tag PathsWebDB 2015
    codedata
  • systemThe ARCOMEM Architecture for Social and Semantic Driven Web ArchivingFuture Internet 2014
    code
  • systemARCOMEM Crawling ArchitectureFuture Internet 2014
    code
  • systemCrawl intelligent et adaptatif d’applications Web pour l’archivage du WebIngénierie des Systèmes d'Information 2014
    codedata
  • demonstrationUne démonstration d’un crawler intelligent pour les applications WebBDA 2013
    code
  • systemCollecte intelligente et adaptative d'applications Web pour l'archivage du WebBDA 2013
    codedata
  • demonstrationDemonstrating Intelligent Crawling and Archiving of Web ApplicationsCIKM 2013
    code
  • systemAn Architecture for Selective Web Harvesting: The Use Case of HeritrixARCOMEM 2013
    code
  • systemIntelligent and Adaptive Crawling of Web Applications for Web ArchivingICWE 2013
    codedata
  • systemSocial and Semantic Driven Web HarvestingBuilding Web Observatories 2013
    code
  • experimental studyOptimizing Approximations of DNF Query Lineage in Probabilistic XMLICDE 2013
    code
  • demonstrationDemonstrating ProApproX 2.0: A Predictive Query Engine for Probabilistic XMLCIKM 2012
    code
  • systemOptimisation des approximations de probabilité des requêtes en XML probabilisteBDA 2012
    code
  • systemExploiting the Social and Semantic Web for guided Web ArchivingTPDL 2012
    code
  • systemAPI Blender: A Uniform Interface to Social Platform APIsWWW 2012
    code
  • systemPARIS: Probabilistic Alignment of Relations, Instances, and SchemaProceedings of the VLDB Endowment 2011
    code
  • systemOntology Alignment at the Instance and Schema LevelBDA 2011
    code
  • demonstrationProApproX: A Lightweight Approximation Query Processor over Probabilistic TreesSIGMOD 2011
    code
  • systemEfficient Query Evaluation over Probabilistic XML with Long-Distance DependenciesJoint EDBT/ICDT Ph.D. Workshop 2011
    code
  • demonstrationUn Système de gestion de données XML probabilistesBDA 2010
    code
  • experimental studyCorroborating Information from Disagreeing ViewsWSDM 2010
    code
  • experimental studyCorroboration de vues discordantes fondées sur la confianceBDA 2009
    code
  • experimental studyAutomatic Wrapper Induction from Hidden-Web Sources with Domain KnowledgeWIDM 2008
    codedata
  • datasetAutomatic discovery of similar wordsSurvey of Text Mining II: Clustering, Classification and Retrieval 2008
    data
  • systemComprendre le Web caché. Understanding the Hidden Web.Miscellaneous 2007
    code
  • experimental studyFinding Related Pages Using Green Measures: An Illustration with WikipediaAAAI 2007
    codedata
  • systemQuerying and Updating Probabilistic Information in XMLEDBT 2006
    code
  • systemQuerying and Updating Probabilistic Information in XMLMiscellaneous 2005
    code
  • experimental studyIdentifying Websites with Flow SimulationICWE 2005
    code
  • experimental studyIdentifying Websites with Flow SimulationMiscellaneous 2005
    code
  • experimental studyWebsite IdentificationMiscellaneous 2003
    code
  • experimental studyAutomatic discovery of similar wordsSurvey of Text Mining: Clustering, Classification and Retrieval 2003
    data
  • systemVérification automatique des multiplicateursMiscellaneous 2002
    code
  • experimental studyExtraction of information in large graphs. Automatic search for synonyms.Miscellaneous 2001
    data

Partly available (5)

  • demonstrationUsing a Probabilistic Database in an Image Retrieval ApplicationEDBT 2025
    codeThe data cannot be redistributed: The degraded COCO images fall under the licenses of the original images; the annotation tables are in the code repository.
  • experimental studyFocused Crawling through Reinforcement LearningICWE 2018
    codeThe data was never distributed.
  • experimental studyRouting an Autonomous Taxi with Reinforcement LearningCIKM 2016
    codeThe data cannot be redistributed: The taxi fleet GPS traces were obtained under a data-sharing agreement.
  • experimental studyHup-Me: Inferring and Reconciling a Timeline of User Activity from Rich Smartphone DataSIGSPATIAL 2015
    codeThe data cannot be redistributed: Personal mobility data.
  • experimental studyTruth Finding with Attribute PartitioningWebDB 2015
    dataThe code was never distributed: The code was kept on a Subversion server that has since closed down.

Nothing available (29)

  • systemAn Indexing Framework for Queries on Probabilistic GraphsACM Transactions on Database Systems 2017
    The code was distributed at the time but has been lost since.
  • experimental studyDiscovering Meta-Paths in Large Heterogeneous Information NetworksWWW 2015
    The code is held by a partner institution: Developed at the Hong Kong University of Science and Technology.
  • demonstrationMonitoring moving objects using uncertain Web dataSIGSPATIAL 2014
    The code was never distributed: The code was kept on a Subversion server that has since closed down.
  • demonstrationCollecte, intégration et visualisation de données Web incertaines sur des objets mobilesBDA 2014
    The code was never distributed: The code was kept on a Subversion server that has since closed down.
  • experimental studyScalable, Generic, and Adaptive Systems for Focused CrawlingHypertext 2014
    The code was distributed at the time but has been lost since: Distributed from a personal Web site that no longer exists.The data was distributed at the time but has been lost since.
  • systemProbTree: A Query-Efficient Representation of Probabilistic GraphsBUDA 2014
    The code was distributed at the time but has been lost since.
  • experimental studyGestion de versions incertaines de documents XMLIngénierie des Systèmes d'Information 2014
    The code was never distributed: The code was kept on a Subversion server that has since closed down.
  • systemContrôle de version incertain dans l'édition collaborative ouverte de documents arborescentsBDA 2013
    The code was never distributed: The code was kept on a Subversion server that has since closed down.
  • experimental studyExploration adaptative de graphes sous contrainte de budgetBDA 2013
    The code was distributed at the time but has been lost since.The data was distributed at the time but has been lost since.
  • experimental studyUncertain Version Control in Open Collaborative Editing of Tree-Structured DocumentsDocEng 2013
    The code was never distributed: The code was kept on a Subversion server that has since closed down.
  • demonstrationCrowd Miner: Mining association rules from the crowdVLDB 2013
    The code is held by a partner institution: Developed at Tel Aviv University.
  • experimental studyCrowd MiningSIGMOD 2013
    The code is held by a partner institution: Developed at Tel Aviv University.The data is held by a partner institution.
  • systemCross-Fertilizing Deep Web Analysis and Ontology EnrichmentVLDS 2012
    The code was distributed at the time but has been lost since.
  • experimental studyOptimal Top-k Generation of Attribute Combinations based on Ranked ListsSIGMOD 2012
    The code is held by a partner institution: Developed at the National University of Singapore.
  • demonstrationAuto-Completion Learning for XMLSIGMOD 2012
    The code is held by a partner institution: Developed at Inria Saclay.
  • demonstrationProFoUnd: Program-analysis–based Form UnderstandingWWW 2012
    The code is held by a partner institution: Developed at the University of Oxford.
  • experimental studyLa diversité culturelle dans l’industrie de la musique enregistrée en France (2003-2008)Publications du Département des études, de la prospective et des statistiques 2011
    The data cannot be redistributed: Proprietary industry statistics.
  • systemXML Content Warehousing: Improving Sociological Studies of Mailing Lists and Web DataBulletin of Sociological Methodology 2011
    The code is held by a partner institution: The aXess platform was distributed by the University of Versailles from a server that no longer exists.
  • demonstrationA Probabilistic XML Merging ToolEDBT 2011
    The code was never distributed: The code was kept on a Subversion server that has since closed down.
  • demonstrationArchivage du contenu éphémère du Web à l’aide des flux WebBDA 2010
    The code was distributed at the time but has been lost since.
  • experimental studyArchiving Data Objects Using Web FeedsIWAW 2010
    The code was distributed at the time but has been lost since.The data was distributed at the time but has been lost since.
  • experimental studyData Quality in Web ArchivingWICOW 2009
    The code is held by a partner institution: Developed at the Max Planck Institute for Informatics.
  • systemThe WebStand ProjectWebSci 2009
    The code is held by a partner institution.
  • experimental studyWeb Page Rank Prediction with Markov ModelsWWW 2008
    The code is held by a partner institution: Developed at the Athens University of Economics and Business.The data is held by a partner institution.
  • systemTraiter des corpus d’information sur le Web. Vers de nouveaux usages informatiques de l’enquête.Miscellaneous 2007
    The code is held by a partner institution: The WebStand platform was distributed by the University of Versailles from a server that no longer exists.
  • systemSYSTRAN Translation Stylesheets: Machine Translation driven by XSLTXML Conference & Exposition 2005
    The code cannot be redistributed: Proprietary SYSTRAN software.
  • experimental studyXML Warehousing Meets SociologyIADIS ICWI 2005
    The code was distributed at the time but has been lost since.
  • systemIntegration of SYSTRAN MT systems in an open workflowMT Summit 2005
    The code cannot be redistributed: Proprietary SYSTRAN software.
  • systemXML Warehousing meets Sociology... Introducing the W3C XQuery Working GroupMiscellaneous 2005
    The code was distributed at the time but has been lost since.

Contact: pierre@senellart.com
  • Everything available (80)
  • Partly available (5)
  • Nothing available (29)

Last Modification
2026-09-03 14:38:01 UTC