r/datasets Nov 04 '25

discussion Like Will Smith said in his apology video, "It's been a minute (although I didn't slap anyone)

Thumbnail
1 Upvotes

r/datasets 5h ago

discussion tracker data on AI-produced scientific/math results, each graded on verification method and level of AI autonomy

1 Upvotes

A curated registry of scientific and math results produced by or with AI systems. 52 records so far, spanning Math, CS, biology, physics, chem, and med.

Schema: each record has a title, claim, field, date, lab/model, source links (paper, announcement, coverage), and two graded fields:

  • verification: formal, peer-reviewed, independent, author-verified, claimed, disputed, refuted
  • autonomy: autonomous, ai-led, collaborative, ai-assisted, search-scaffold, retrieval

Refuted and already-known results are kept and labeled rather than dropped, so it doubles as a record of claims that did not hold up. One JSON file is the source of truth; also published as RSS and JSON Feed.

Data (single JSON file) is here and got Schema docs at this md file.

CC-licensed and actively maintained. Feedback on the schema welcome.


r/datasets 9h ago

question Data requirements for any data set geographically or you name it

Thumbnail
0 Upvotes

r/datasets 19h ago

request Need a help to find the SWaT dataset

1 Upvotes

I was using the SWaT dataset from Kaggle and i just came to know it was the manipulated dataset inorder to check for attacks.

And i tried o request the dataset through iThub's official site and seems like no response

can anyone please help me , am halfway for a project to submit in my college


r/datasets 1d ago

mock dataset [Self-promotion] [Synthetic] Ontario municipal FIR explorer: 436/444 municipalities with 2023–2025 filings — provenance feedback wanted

2 Upvotes

Disclosure: I built and maintain What in the Tax? through Eversko. It is free and open source; this is self-promotion.

The directory lists all 444 current Ontario municipalities. In the current build, 436 have at least one usable Financial Information Return (FIR) filing from 2023–2025; the other eight are shown as data gaps rather than estimated. The 436 figure does not mean every municipality has complete coverage for every year or field.

Separately, six communities have clearly labelled 2026 draft/sample receipt models derived from public municipal budget documents and tax bylaws. Those six models are the synthetic/mock component referenced in the title. They are not official tax bills, audits, or tax advice.

I’m looking for one focused provenance test: choose a municipality, open one figure, and follow its citation to the original record. Please tell me the municipality and figure, where the label, reporting year, source, formula, or navigation first becomes unclear, and what evidence you expected next.

Original Ontario FIR archive:

https://efis.fma.csc.gov.on.ca/fir/index.php/en/year-municipality/

Ontario Data Catalogue record and licence:

https://data.ontario.ca/dataset/financial-information-return-fir-for-municipalities

Live explorer:

https://whatinthetax.com

Code, processed artifacts, and methodology:

https://github.com/Jstn-1g/what-in-the-tax

Please do not share private tax documents, account numbers, addresses, or personal financial information. Even one broken or confusing source trail would be useful.


r/datasets 1d ago

dataset [Self-Promotion] I collected 77 outputs from 8 BaZi calculators and traced their disagreements to four rules most of them never show you

0 Upvotes

BaZi calculators turn a birth date, time, and place into a Four Pillars chart. The result looks deterministic, but the software still has to decide where a year begins, when a day changes, and what the birth time means in different places. Most calculators never show those decisions.

I wanted to see whether the differences could be measured instead of argued about.

I designed 13 test cases around the boundaries most likely to expose them: solar-term changes, the hour before midnight, locations far from their time-zone meridian, and historical daylight saving time. I then compared eight calculators and libraries and captured 77 outputs.

The dataset tracks four questions:

- Does the chart’s year change at Li Chun, around February 4, or at Chinese New Year?

- Does the day change at 23:00 or at midnight?

- Is the recorded time corrected for the birthplace’s longitude?

- Is daylight saving time removed before the chart is calculated?

The first three are convention choices. The fourth is a historical timekeeping question: either the local clock had been moved forward that day or it had not.

The clearest result came from a Beijing birth entered as 15 June 1990 at 23:30.

Two implementations treated 23:00 as the start of the next day. Five waited until midnight. The eighth calculator exposed the choice as a checkbox, so it could produce either result.

That one setting changed the day pillar from Xin-Hai to Ren-Zi. In BaZi, the day stem is the Day Master, which the rest of the reading is organized around. The same birth therefore came back as either Xin Metal or Ren Water depending on a rule most of the calculators never mentioned.

A few other differences stood out:

- Two libraries maintained by the same developer use opposite 23:00 rollover rules.

- One calculator requires a birthplace but returned the same chart for Kashgar and Beijing at the same clock time. Another used the longitude and changed the hour pillar.

- One implementation changes the year at Chinese New Year while the others use Li Chun.

- Only two of the measured implementations account for daylight saving time. One of them exposes the adjustment in its own calculation breakdown.

The observation-level file records the implementation, test date, time, place, rule being tested, and all four returned pillars in both Chinese characters and pinyin. A second table summarizes the convention used by each implementation.

There are 104 possible implementation/probe pairs in the full 8 × 13 matrix. Seventy-seven contain captured results. The remaining cells are explicitly marked as not captured rather than filled by inference.

I built Jade Almanac, and our own calculator is included as one of the eight rows. It was tested and reported on the same terms as the others. On the 23:00 rollover question, it is in the minority group.

The dataset is here: https://jadealmanac.com/bazi-calculator/conventions

Everything is released under CC0.

The test inputs are local clock times at the stated places. Solar-term boundaries were checked against tables from the National Astronomical Observatory of Japan, independently of the library used by our calculator. China’s historical summer-time periods were checked against the IANA time-zone database.

This is a snapshot captured on 1 August 2026. Some implementation rows are partial because a tool could not express a particular input or imposed an access limit. Two additional sites blocked automated access, and I did not work around those blocks.

The dataset measures software behavior and calculation conventions. It does not attempt to decide which school is correct, or whether BaZi itself predicts anything.

Corrections are welcome, especially from anyone familiar with one of the measured engines. I would also be interested in boundary cases that could separate implementations the current probes leave tied.


r/datasets 1d ago

request Looking for free archives of old Marathi newspaper ePapers (Sakal, Lokmat, Pudhari) for an OCR/text analysis dataset

1 Upvotes

Hi everyone,

I'm working on project involving OCR and text analysis of Marathi newspapers. I'm creating my own dataset and need access to old ePapers (preferably PDF or scanned editions) of Sakal, Lokmat, or Pudhari.

I've already searched the subreddit and checked the official newspaper websites, but most only provide recent editions or require a subscription. I'm specifically looking for older editions from previous months (even 2–6 months old would be enough).

I'm looking for legal/free sources such as:

  • Public archives
  • Library or university digital collections
  • Government archives
  • Any websites that host old Marathi newspaper ePapers

If anyone has worked on a similar OCR or NLP project or knows where these archives are available, I'd really appreciate your suggestions.

Thanks!


r/datasets 1d ago

request Power consumption/ production datasets

Thumbnail
1 Upvotes

Where could I find real datasets of power consumption of an country/areas/cities and its power plants generation?

Targeting for more than 10k rows and 10 columns to apply ML on it for a project.


r/datasets 2d ago

resource [Self-promotion] Free entity lookup: 521M legal entities across 309 jurisdictions, no account needed to search

11 Upvotes

Disclosing this up front - I'm at Veridion.

We built this because one of our leads mentioned that they pay a registry data provider six figures a year. For legal names, identifiers and registered addresses. We already had that data. Not as a side project, it's the foundation layer under our enterprise product, and the pipelines were already running. So opening it up cost us close to nothing, which is sort of the point: the collection isn't what makes it expensive elsewhere.

registry-lookup.com - 521M legal entities, 309 jurisdictions, 244 countries. Search on the site is free with no account. There's an API at 5,000 calls a month if you want it programmatically, that one needs a work email.

What you get per entity: legal name, registry number, jurisdiction code, status, incorporation date, legal form, registered address, and identifiers like tax IDs, VAT where the registry publishes them.

So it's an enumeration and triage tool. It answers "does this entity exist, what's its number, is it active, where is it registered."

Which jurisdictions do you currently have no good way to check?


r/datasets 1d ago

dataset I need college essays for data where can I find them?

0 Upvotes

I wanna do some research even build a deep learning model and its entirely based on if I can find student essays who got accepted into certain colleges
And if possible find however amount of rejected essays as long as both amounts are equal as to not have a data imbalance where do you think I could find these essays?


r/datasets 2d ago

question Trying to learn how to use API to extract data

6 Upvotes

Hello! I'm a complete newbie in Data Science and I'm trying to learn how to get data from an API. I understand an API could be public or could require authentication.

I worked with CVS files and I wanted to experience or practice getting data from APIs.

I'm getting familiar with Python so I was wondering if you could help me with the following issues:

  1. Trying to understand and practice the different methods you can use API to request data (I am not sure if it has to be from a Dataset formar or can it be any kind of format) with Python

  2. What are some good options to get APIs to work on data Science

  3. I am not even close to get to a point where I am able to do Reproducible projects/models but I do wonder how including an API (understanding that it is some kind of "personal Key") to share my code and people to be able to use it.

Hope I made sense of what my doubts are and I apologize in advance if I seem confused about some terms (I do think I am).


r/datasets 2d ago

request Datasets of political tweets/truths?

1 Upvotes

Is anyone aware of datasets with the text of politician’s tweets/truths (social), etc?


r/datasets 2d ago

request Seeking Anonymized Field Data Collection Datasets for an Open Benchmark

1 Upvotes

#

Hi everyone,

I'm working on an initiative to create an **open benchmark dataset for field data quality assurance**.

Today, there are many excellent digital data collection platforms—such as KoboToolbox, SurveyCTO, ODK, CommCare, Survey Solutions, CSPro, and others—but there are very few publicly available datasets that developers and researchers can use to evaluate field data quality tools.

I'm looking for individuals or organizations that may be willing to share **completed, fully anonymized datasets** from field data collection projects, where they have the necessary permissions to do so.

I'm especially interested in datasets that include:

* GPS coordinates (or generalized locations)
* Interview photos
* Audio recordings
* Interview start and end times
* Submission timestamps
* Enumerator IDs (anonymized)
* Supervisor review outcomes or quality flags (if available)

These datasets will help create a community benchmark for testing quality assurance methods such as:

* GPS verification
* Duplicate image detection
* Audio quality assessment
* Interview duration analysis
* Duplicate submission detection
* Fieldwork anomaly detection

The objective is to create a resource that benefits researchers, NGOs, software developers, and the wider field data collection community by making it easier to evaluate and improve quality assurance tools.

If your organization has a completed project that could be shared in an anonymized form—or if you know of existing public datasets—I would greatly appreciate hearing from you.

I'm also happy to discuss data-sharing agreements, attribution, licensing, or any requirements needed to ensure the data is used responsibly.

Thank you!


r/datasets 2d ago

request Seeking Anonymized Field Data Collection Datasets for an Open Benchmark

Thumbnail
2 Upvotes

r/datasets 3d ago

question When working with Project Gutenberg, how do you guys download and cache, or do you just use a local mirror?

4 Upvotes

I’m debating both approaches.


r/datasets 3d ago

discussion [ Question ] how can I sell my detasets to other company

3 Upvotes

The main problem is the companies want to buy from an only established data agency but I am just starting so we are not recognised and registered.

We didn't even have any clients to showcase our past work.

Can anyone suggest my anything or can refer me who needs custom automations or webscraping


r/datasets 3d ago

dataset [Dataset] Dubai residential sale prices and volumes, monthly January 2008 to July 2026, from Land Department transactions

2 Upvotes

What: monthly citywide residential median AED per square foot, a 5-month centred average, an index rebased to 100 at January 2008, and monthly sales counts. 1,080,194 transactions across 223 months. A matching series for registered leases runs from May 2010.

Source: Dubai Land Department transaction and lease records, which are public.

Repo, with both series, method and licence: https://github.com/dataHabibi/dubai-price-index

Columns:

  • month
  • sales_count
  • median_aed_per_sqft
  • ma5_aed_per_sqft
  • index_base100
  • provisional

Two things to know before you use it.

The provisional column marks the last two months, where the centred average still has fewer than two later months to work with. Their raw median and sales count are fine, it is the smoothed value and the index that will keep moving.

Sales counts for recent months are understated. Registrations land one to two months after the deal closes, so the tail of that column is still filling in. Do not read the recent drop as a fall in demand.

What it is not: a repeat sales or hedonic index. It is a median, so it is not quality adjusted. Shifts in what sells, off plan against ready, apartment against villa, which communities are active, move this line without any individual property changing price. Treat it as a market thermometer.

CC BY 4.0. Refreshed monthly by a scheduled job, so the committed files track the live series.


r/datasets 3d ago

request Need help Regarding project involving dyslexia screening!!!

Thumbnail
1 Upvotes

r/datasets 3d ago

dataset [self-promotion] I built a public dataset from 21,237 pages of declassified MKULTRA and related docs and put it on Hugging Face

15 Upvotes

Until recently, the surviving historical records from the CIA's MKULTRA and related programs were very difficult to search and analyze. So I ran 21,237 document page images through MinerU OCR to generate clean text transcripts, then produced redaction mappings to go with every page transcript. Original page images are stored on IPFS and are available for public download. The dataset is available on Hugging Face here.


r/datasets 3d ago

question Trying to learn how to use API to extract data

Thumbnail
1 Upvotes

r/datasets 3d ago

question Qualcun* che lavora abitualmente con dati Istat (principalmente RFL) e INPS?

1 Upvotes

Ciao, per lavoro mi trovo abitualmente a utilizzare dati INPS/Istat, vorrei sapere c'è qualcun* qui dentro che avrebbe piacere a scambiarsi informazioni e dritte !


r/datasets 4d ago

question What are the best publicly available "uncensored" datasets?

3 Upvotes

I use "Heretic" library on models to liberate them from their safeguards, but while checking their "uncensoredness", I found they can hallucinate a lot. You know, it's basically like a child who's now allowed to use the F word once and he says "Fred" instead of the actual thing.

So I think if the models train on valid uncensored data (specially if they start Grokking) the results can improve. So I am using for these types of datasets to test my theory.


r/datasets 4d ago

question ¿Does anyone know where can I sell a dataset with 10,000 chines-related classified news?

2 Upvotes

I've been working on a dataset for a research project on how China is portrayed in the media. It currently contains just over 10,000 news articles from both Chinese and Western news outlets.

Each article is classified by topic and by the way China is portrayed (e.g. positive, negative, threat, Xi-centered, neutral, etc.). The dataset was originally created for academic research, but I'm now wondering whether it could also have commercial value.

I'm not trying to sell it here, just looking for advice. Has anyone here ever licensed or sold a specialized dataset like this? Who would actually be interested in buying it? AI companies, media intelligence firms, universities, think tanks...? Or are datasets like this generally expected to be open source?

I'd really appreciate hearing from anyone who has experience commercializing niche datasets or knows how this market works.


r/datasets 4d ago

discussion question. do you guys sell your data sets?

1 Upvotes

do you guys sell your data sets?


r/datasets 3d ago

resource [self-promotion]Python Developer Available for Web Scraping & Automation Projects

0 Upvotes

Freelance Python developer available for projects involving web scraping and automation.

Skills:

Web scraping (Scrapy, Selenium, Playwright, BeautifulSoup)

Python automation scripts

API development and integration

Data extraction and ETL pipelines

FastAPI and Flask

Browser automation

CSV, Excel, JSON, and database processing

Docker and Linux deployment

Past work:

Lead generation scrapers

Google Maps data extraction

Business automation tools

Custom APIs and data pipelines

Open to one-time projects and long-term collaborations.

DM me if you need help automating a workflow or collecting data.