Builds a graph database of Instagram accounts.
This project was created by the CBS News Data Team to map networks of accounts that post violent content on Instagram Reels. It aims to provide insights into how such accounts interact and influence each other, although it can be used to map any network of Instagram accounts.
The project is written in Python and consists of:
- A web crawler built using Playwright
- Tools to fine-tune SetFit models to prioritize accounts for crawling
- A method to load the crawl results into a Memgraph database for analysis.
Crawling Instagram without being blocked is quite difficult, so this project takes the approach of crawling very slowly and carefully to avoid detection. It also uses a SetFit model to prioritize accounts that are more likely to be problematic based on their account names. In our experience, this approach has been successful in crawling thousands of accounts without being blocked, although it can take days or longer to crawl a large network.
To set up instagraph in a basic python environment, install the package from Github:
pip install git+git@github.com:cbs-news-data/instagram-reels-violence.gitYou'll also need to run Memgraph in order to load the crawl results into a graph database. See the Memgraph documentation for instructions on how to set it up.
The recommended way to run the crawler is using Docker and docker-compose. To set up the project using docker, you'll need to create a docker-compose file in your project directory. That docker-compose file must define several services:
- The crawler service, which will run the crawler
- The Memgraph service, which will run the Memgraph database
- (Optional) The Memgraph Lab service, which will run the Memgraph Lab interface for exploring the database
In addition, you can run multiple crawler instances simultaneously by defining multiple crawler services in the docker-compose file. You can also run a service that continuously loads crawl results into the database by defining a service that runs the load command.
See our instagram-reels-violence repository for a full example of how to set up the project using Docker.
Viewing account data like this requires authenticating the Crawler with a valid Instagram account. The next step is to get a pool of Instagram accounts for the crawler to use. How you create these accounts is up to you, but I recommend using a service like simplelogin to create a pool of email addresses that can be used to create Instagram accounts. You'll need to manually create these accounts and provide the credentials to the crawler.
The crawler is highly configurable and can use several strategies to determine which accounts to crawl. To view the full list of options, run:
instagraph crawl --helpThe most basic usage of the crawler is to provide a list of seed accounts to start from. I recommend starting with a small number of accounts and doing a depth-first search on the first run. For example:
instagraph crawl <problematic_account_1> <problematic_account_2> <problematic_account_3> --enqueue-accounts --enqueue-depth 5This will crawl the accounts provided and enqueue the accounts they follow. The --enqueue-accounts flag tells the crawler to continue crawling accounts followed or mentioned in bio by the seed accounts, and the --enqueue-depth flag specifies how many levels deep to crawl. The crawler will prioritize accounts that are more likely to be problematic based on the SetFit model, if you've fine-tuned one using the fine-tune command (see the Fine tuning section for details).
Once you've run a crawl, you can either continue to manually provide additional accounts, or use the --resume-queue flag to automatically enqueue new accounts that are discovered during previous crawls, but weren't crawled themselves. For example:
instagraph crawl --resume-queue --enqueue-accountsThis will continue to crawl new accounts until there are no more accounts to crawl.
Once you've started the Memgraph database, you can use the --start-query-path flag to start the crawl with the results of a query. For example:
instagraph crawl --start-query-path queries/seed_accounts.cypher --reels-limit 100 --enqueue-accounts --enqueue-depth 5You can provide any query here, as long as each row it returns contains a column called "account_name," which will be used as the seed accounts for the crawl. The best way to use this is to run a query to find un-crawled nodes in the graph, and use those as the seed accounts for the next crawl.
For example, this query finds the most central accounts in the graph that are related to accounts that have sensitive content, but haven't been crawled yet:
MATCH p=(s:Account {has_sensitive_content:"true"})-[r:FOLLOWING|:MENTIONS_IN_BIO]-(a:Account {date_crawled:""})
WITH p, a
WITH project(p) as subgraph
CALL katz_centrality.get(subgraph)
YIELD node, rank
WITH node, rank
WHERE rank >= 0.33 AND node.prediction = "positive"
RETURN node.account_name AS account_name, rank
ORDER BY rank DESC;
The crawler can be run in a variety of ways, depending on your needs. You can run it manually, or schedule it to run at regular intervals. The crawl and load commands use the schedule library to schedule tasks if you provide the --every flag. Run the following command to see the full list of options:
instagraph crawl --helpor
instagraph load --helpFor example, to schedule a crawl to run every 6 hours, run:
instagraph crawl <other-args-to-start-crawl> --every 6 --unit hoursIf you wanted to re-crawl a single account and all the accounts it follows every 24 hours, you could run:
instagraph crawl <account> --enqueue-accounts --enqueue-depth 1 --every 24 --unit hoursWhile you have that crawler running, you can also run a separate service to load the crawl results into the Memgraph database. For example, to load the crawl results into the database every 10 minutes, run:
instagraph load --every 10 --unit minutesYou can provide those commands in a docker-compose file to run them in separate containers for a more robust setup.
The crawler uses a SetFit model to prioritize accounts that are more likely to be problematic based on their account names. The fine-tune command will help you interactively fine-tune a SetFit model using a list of account names. For example:
instagram fine-tuneNOTE: this command requires you to have crawled some accounts first, so that you have a list of account names to fine-tune the model with. If you haven't crawled enough accounts, it will throw an error.
Once you've fine-tuned a model, the crawler will automatically use it to prioritize accounts for crawling. If you re-fine-tune the model by running the fine-tune command again, the crawler will automatically use the newest model.
Once you've crawled some accounts, you can load the data into Memgraph for analysis. The load command will load the crawl results into a Memgraph database. For example:
instagraph loadThis will create a graph database in Memgraph, which you can then query using the Memgraph query language.