A system called 'RSL' will be developed to notify terms of use and fees for scraping for AI learning purposes, and Yahoo, Reddit, O'Reilly, and others have already announced their adoption

Developing AI requires massive amounts of data, and AI development companies use automated bots (scrapers) to collect all kinds of information available on the internet. A system called ' Really Simple Licensing (RSL) ' has been developed that allows users to set terms of use and fees for these scrapers. The RSS developers and Tim O'Reilly, founder of O'Reilly Media, were involved in the development, and services such as Yahoo, Reddit, O'Reilly Media, Quora, and Medium have already announced their adoption.
RSL
New RSL Web Standard and Collective Rights Organization Automate Content Licensing for the AI-First Internet and enable Fair Compensation for Millions of Publishers and Creators | RSL
https://rslstandard.org/press/rsl-standard
Web developers can use RSLs to communicate policies to scrapers, such as 'prohibit use for AI training,' 'allow use for AI training without restrictions,' or 'allow use for AI training if a usage fee is paid.' Incorporating RSLs into a website is simple: simply place a 'license.xml' file containing the license terms in the root directory and add the location of license.xml to robots.txt. Detailed instructions and license terms templates are available at the following link.
Getting Started | RSL
https://rslstandard.org/guide/getting-started

RSLs can also be applied to individual pages within a website , and by adding RSLs to an RSS feed, the RSS feed can be treated as a 'standardized catalog of licensable digital assets.'
Adding RSL to RSS feeds | RSL
https://rslstandard.org/guide/rss-feeds

The RSL steering committee includes RSS co-creators Eckart Walther and Ramanathan V. Guha, as well as O'Reilly Media founder Tim O'Reilly and Fastly co-creator Simon Wistow.

Additionally, publishers such as Reddit, People Inc., Yahoo, Internet Brands, Ziff Davis, wikiHow, O'Reilly Media, Medium, The Daily Beast, Miso.AI, Raptive, Ranker, and Evolve Media have already announced their adoption of RSL.

However, RSL is merely a function to notify scrapers of license terms, and it is unknown whether scrapers will actually respect the license terms. In fact, there have been reported cases where AI companies have ignored websites' instructions not to crawl and collected information.
Cloudflare accuses Perplexity of using stealth tactics to ignore no-crawl orders - GIGAZINE

Related Posts:
in AI, Web Service, Security, Posted by log1o_hf







