scraping-tweets-by-keyword

v2026.09.24

Scrapes tweets matching any keyword, hashtag, phrase, or boolean query using apidojo's Twitter Search scraper on Apify. Triggers when the user asks to: scrape tweets about a topic, fetch tweets containing a keyword or hashtag, export Twitter search results to a dataset, get all tweets mentioning a phrase, pull recent tweets for a search term, or collect tweet data for analysis or research. Returns tweet text, author, likes, retweets, replies, timestamp, and tweet URL per result. Ideal for researchers, data analysts, journalists, and social listening teams.

GitHub
Install command
npx skhub add apidojo-io/scraping-tweets-by-keyword
Markdown
SKILL.md

Scraping Tweets by Keyword

Raw tweet collection for any keyword, hashtag, or boolean search query. No assumed use case — returns the full tweet dataset for downstream analysis.

Prerequisites

  • APIFY_TOKEN environment variable set
  • Optional: Apify MCP server installed

Inputs

ParameterTypeRequiredDefaultNotes
searchTermsarray✅[]Twitter advanced search queries (e.g. ["#AI lang:en", "from:NASA"])
sortstringOptionalTopSort order: Latest, Top, or Latest+Top
tweetLanguagestringOptional—ISO 639-1 language code (e.g. en)
maxItemsnumberOptionalUnlimitedMaximum tweets to return
onlyVerifiedUsersbooleanOptionalfalseOnly tweets from verified users
onlyTwitterBluebooleanOptionalfalseOnly Twitter Blue subscribers
onlyImagebooleanOptionalfalseOnly tweets with images
onlyVideobooleanOptionalfalseOnly tweets with videos
onlyQuotebooleanOptionalfalseOnly quote tweets
authorstringOptional—Filter to a specific author handle
inReplyTostringOptional—Tweets replying to a specific handle
mentioningstringOptional—Tweets mentioning a specific handle
geotaggedNearstringOptional—Tweets near a location
withinRadiusstringOptional—Radius around geotaggedNear
geocodestringOptional—Lat/lng + radius string
placeObjectIdstringOptional—Tweets tagged with a place
minimumRetweetsnumberOptional—Minimum retweet count
minimumFavoritesnumberOptional—Minimum like count
minimumRepliesnumberOptional—Minimum reply count
startstringOptional—Tweets after this date (YYYY-MM-DD)
endstringOptional—Tweets before this date (YYYY-MM-DD)
includeSearchTermsbooleanOptionalfalseAdd the matched search term to each tweet
customMapFunctionstringOptional—JavaScript function to transform each output object

Workflow

Progress:
- [ ] Step 1: Build search query string
- [ ] Step 2: Run tweet-scraper
- [ ] Step 3: Poll for SUCCEEDED
- [ ] Step 4: Fetch and deliver dataset

Step 1: Build Search Query

  • Hashtag search → #keyword
  • Exact phrase → "exact phrase"
  • Boolean → word1 AND word2 -exclude
  • From account → from:username
  • Mention → @username

Step 2: Run the Actor

Recommended — run_actor.js (handles waiting, output, and file saving automatically):

# Quick answer (prints table to chat)
node scripts/run_actor.js \
  --actor "apidojo~tweet-scraper" \
  --input '{"param": "value"}'

# Save as CSV
node scripts/run_actor.js \
  --actor "apidojo~tweet-scraper" \
  --input '{"param": "value"}' \
  --output YYYY-MM-DD_results.csv --format csv

# Save as JSON
node scripts/run_actor.js \
  --actor "apidojo~tweet-scraper" \
  --input '{"param": "value"}' \
  --output YYYY-MM-DD_results.json --format json

APIFY_TOKEN must be set in environment or .env file.

If Apify MCP is available:

Tool: apify:run-actor
Actor: "apidojo~tweet-scraper"
Input:
{
  "searchTerms": ["<query>"],
  "maxItems": 200,
  "since": "<YYYY-MM-DD>",
  "lang": "<lang_code>"
}

REST API fallback:

curl -X POST \
  "https://api.apify.com/v2/acts/apidojo~tweet-scraper/runs?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"searchTerms": ["<query>"], "maxItems": 200}'

Save id as RUN_ID. Poll until status = SUCCEEDED:

curl "https://api.apify.com/v2/actor-runs/$RUN_ID?token=$APIFY_TOKEN" | grep '"status"'

Fetch results:

curl "https://api.apify.com/v2/actor-runs/$RUN_ID/dataset/items?token=$APIFY_TOKEN&format=json"

Step 3: Handle Edge Cases

  • 0 results: Query may be too narrow, misspelled, or language-filtered. Broaden term, remove language filter, extend date range.
  • < 20 results: Try removing since/until constraints. Some low-volume terms have sparse data.
  • Duplicate tweet IDs: Deduplicate by id field before delivering.
  • Suspended/deleted accounts: Tweets from suspended accounts return with empty author fields — flag these rows.

Output Format

# Tweet Dataset: "<query>"
Total collected: N | Date range: SINCE – UNTIL | Language: LANG

| Tweet ID | Author | Text (truncated) | Likes | Retweets | Replies | Timestamp |
|----------|--------|-----------------|-------|----------|---------|-----------|
| ...      | ...    | ...             | ...   | ...      | ...     | ...       |

Full dataset: N rows × 15 fields
Available fields: id, text, author_id, author_username, likeCount, retweetCount,
replyCount, quoteCount, lang, createdAt, tweetUrl, media, isRetweet, isQuote, source

Troubleshooting

Empty results for a valid hashtag: Twitter API indexing lag — try again after 15 minutes. Rate limit error: Reduce maxItems to 100 and retry. Timeout on large requests: Set maxItems: 500 max per run; chain multiple runs with date ranges for larger datasets.

Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

Apache-2.0

Source path

skills/primitives/scraping-tweets-by-keyword

Default branch

main

Latest commit

ffbdc00

Tree SHA

c7df562