University of Chicago research group by
Ciro Santilli 37 Updated 2025-05-29 +Created 2025-02-26
This is the family of algorithms to which PageRank
Open PageRank implementation and data by
Ciro Santilli 37 Updated 2025-05-29 +Created 2025-02-26
This section is about more "open" PageRank implementations, notably using either or both of:
As of 2025, the most open and reproducible implementation appears to be whatever Common Crawl web graph official PageRank does, which is to use WebGraph. It's quite beautiful.
École Polytechnique alumnus of 2009 by
Ciro Santilli 37 Updated 2025-05-29 +Created 2025-02-26
École Polytechnique alumnus of 1983 by
Ciro Santilli 37 Updated 2025-05-29 +Created 2025-02-26
In 2017 apparently they've started making their own Web Graphs, i.e. they parse the HTML and extract the graph of what links to what.
Edit: actually, they already calculate PageRank for us!!! Fantastic!!! Main section: Section "Common Crawl web graph official PageRank".
A quick exploration of the graph can be seen at: github.com/cirosantilli/cirosantilli.github.io/issues/198
Their source code is at: github.com/commoncrawl/cc-webgraph
École Polytechnique alumnus by year by
Ciro Santilli 37 Updated 2025-05-29 +Created 2025-02-26
École Polytechnique students identify their academic year, or "promotion" in French, by start year date.
For example, Ciro Santilli's year started in 2009, though as a foreign student he arrived only at the start of 2010, and Ciro's promotion is usually known just as X09. And as the century barrier is broken we'll start to need to specify as X2009 one day.
List of notable alumni:
A quick hands-on introduction to the software by Ciro Santilli can be found at: github.com/cirosantilli/cirosantilli.github.io/issues/198
The native file format of WebGraph.
It is a binary format and highly storage efficient.
It is for example what Common Crawl web graph currently dumps to as of 2025, see e.g.: data.commoncrawl.org/projects/hyperlinkgraph/cc-main-2024-25-dec-jan-feb/index.html
TODO meaning of "BV"?
A quick hands-on introduction to the format by Ciro Santilli can be found at: github.com/cirosantilli/cirosantilli.github.io/issues/198
Unlisted articles are being shown, click here to show only listed articles.