Alice Magnusson

Alice Magnusson

Data Teams, Developers, Operators

About

Research data engineer in a university lab. I work on data pipelines, reproducible analysis, dataset versioning, provenance, and the unglamorous parts of research that keep results trustworthy.

Badges

Tastemaker
Tastemaker
Gone streaking
Gone streaking

Forums

Who is the best proxy provider for web scraping?

Over the last four years, we've used, tested, and benchmarked close to 100 web scraping APIs and residential proxy providers across different websites and request volumes.

One thing has become clear to us: there is no single best proxy provider.

How Much Does It Cost to Generate Web Scraping Code With AI?

AI is starting to change the economics of building scrapers. A scraper that used to take a few hours of manual selector work, browser debugging, schema cleanup, retries, and maintenance can now often be scaffolded in minutes.

The part I'm interested in is the real cost, not the demo cost.

3mo ago

Working on reproducible research data

Hi, I'm a research data engineer working with academic teams on data collection, cleaning, versioning, and reproducible analysis.

A lot of research data work breaks down before the model, chart, or paper ever happens. Sources change, scripts are undocumented, licenses are unclear, and six months later nobody can explain exactly how the dataset was created.

I'm especially interested in open science, dataset provenance, archival workflows, and practical ways to make research easier to rerun without slowing everyone down too much.

Curious how others here handle this: when you publish or share a data project, what do you document first... raw data, collection timestamps, cleaning scripts, licenses, environment files, or something else?

View more