University of Milan Updated +Created
PhD thesis Updated +Created
WebGraph (software) Updated +Created
BVGraph Updated +Created
The native file format of WebGraph.
It is a binary format and highly storage efficient.
TODO meaning of "BV"?
A quick hands-on introduction to the format by Ciro Santilli can be found at: github.com/cirosantilli/cirosantilli.github.io/issues/198
Cancer research Updated +Created
Updates / Quick fun with the Common Crawl web graph Updated +Created
I wanted to do a quick exploration of open PageRank implementation and data.
My general motivation for this is that a PageRank-like algorithm could be useful for more accurate user and article ranking on OurBigBook, see: Section "PageRank-like ranking"
But it could also be just generally cool to apply it to other graph datasets, e.g. for computing an Wikipedia internal PageRank.
A quick Google reveals only Open PageRank, but their methods are apparently closed source.
Then I had a look at the Common Crawl web graph data to see if I could easily calculate it myself, and... they already have it! See: Section "Common Crawl web graph official PageRank"
Their graph dumps are in BVGraph graph file format, which is the native format of the WebGraph framework, which implements the format and algorithms such as PageRank.
The only thing I miss is a command line interface to calculate the PageRank. That would be so awesome.
The more I look at it the more I love Common Crawl.
In cc-main-2024-25-dec-jan-feb-domain-ranks.txt:
  • cirosantilli.com was ranked ~453k
  • ourbigbook.com was at ~606k
White-East asian mixed Updated +Created
Multiracial Updated +Created
White people Updated +Created
Race (human categorization) Updated +Created
Atherton, California Updated +Created
Municipality in San Mateo County Updated +Created
San Mateo County Updated +Created
Municipality in Illinois Updated +Created
County in California Updated +Created
Flag of the Soviet Union Updated +Created
East Asian Updated +Created
Eastern Europe Updated +Created
Apple Verdiell Updated +Created

Unlisted articles are being shown, click here to show only listed articles.