Using deep learning to detect abusive sequences of member activity
Originally published Updated 4 min read
Detecting logged-in profile scrapers from the order and timing of their requests.
One logged-in profile scraper we studied at LinkedIn viewed roughly seventy distinct profiles in a day with randomized delays. Another viewed profiles in short bursts and deliberately revisited profiles it had already seen. Their activity volumes could plausibly have been human; request counts alone gave us little separation.
Beibei Wang and I developed a production deep learning model to detect profile scrapers from the order and timing of their requests. We described the model in our LinkedIn Engineering article.
Problem and input data
Request sequences can help identify automation used for fake accounts, account takeovers, API abuse, and scraping. Our first production application was logged-in profile scraping.
From Requests to Tokens
We modeled the ordered requests from an account, including the requests around each profile view: logins, searches, messages, settings changes, and page resources.

A mock burst of profile views and the request sequence surrounding it. The time gaps between requests are supplied to the model separately. (Figure from my LinkedIn Engineering blog post.)
The pipeline canonicalizes raw request paths into standardized path tokens, then assigns integer IDs in global request-frequency order: common requests get small IDs and rare requests get large IDs. These IDs index learned request-path embeddings.
The standardized request paths form a sequence of tokens, analogous to the words in a sentence. The elapsed time between each pair of requests is supplied as a second input.
What the Model Sees
The grid shows 200 consecutive requests, twenty per row, colored by how common each request is. Here is a legitimate member browsing the site:

Request-path tokens for 200 consecutive requests from a legitimate member, colored by request frequency. (Figure 2 from our LinkedIn Engineering article.)
And here is a scraper:

The same visualization for this scraper is dominated by common request types, with little variation. (Figure 3 from our LinkedIn Engineering article.)
The grids show request order and type; the separate timing input is not shown.
Model architecture
The classifier has separate request-path and timing branches. The request-path branch begins with embeddings learned during supervised training. One-dimensional convolutions identify short subsequences within the embedded request stream.
The timing branch processes the inter-request time gaps. After the path and timing representations are concatenated, an LSTM models longer-range dependencies across the account’s activity window. A final dense layer produces an abuse score.

Request-path and timing branches of the sequence classifier. (Figure 4 from our LinkedIn Engineering article.)
Training labels
We used the production Isolation Forest model to generate weak labels for the sequence classifier. Examples from known historical attacks could also be added to the training data.
Evaluation at Natural Class Balance
We evaluated the initial proof-of-concept model on data collected well after the training period, retaining the natural class balance.

Score distributions for the out-of-time test set. The scraper labels came from Isolation Forest; the figure also distinguishes accounts with high Isolation Forest scores. (Slide 34 from my Scale AI talk.)
In the highlighted high-score bins, accounts labeled as scrapers outnumbered accounts labeled as non-scrapers by roughly 1,000 to 1.
Activity sequence embeddings
The model also produces activity sequence embeddings that can be used to group accounts with similar request patterns for investigation.

A two-dimensional projection of activity sequence embeddings, colored by the non-scraper and scraper label groups. (Slide from my Scale AI talk.)
The embeddings can also serve as features for downstream outlier-detection models, replacing hand-engineered activity summaries with learned representations of the sequence.
Resources
- Using deep learning to detect abusive sequences of member activity (LinkedIn Engineering)
- Talk recording (YouTube)
- US Patent 11,936,682: DEEP LEARNING TO DETECT ABUSIVE SEQUENCES OF USER ACTIVITY IN ONLINE NETWORK
- US Patent 11,991,197: DEEP LEARNING USING ACTIVITY GRAPH TO DETECT ABUSIVE USER ACTIVITY IN ONLINE NETWORKS
- US Patent 12,500,923: IDENTIFYING COORDINATED MALICIOUS ACTIVITIES USING SEQUENCES OF REQUESTS