【问题标题】:Twitter streaming with multiple twitts that have same id具有相同 id 的多个 twitts 的 Twitter 流
【发布时间】:2015-04-30 09:14:17
【问题描述】:

我在this pipeline 的帮助下收集推文。我尝试使用一些自己的脚本来分析收集的脚本。我发现我收到了多条具有相同 ID 的推文。我查看了 hdfs://user/flume/tweets 并看到这多条推文在存储的文件中。所以这不是 hive 或 oozie 问题。

可能是水槽问题:我对水槽参数进行了一些编辑:

TwitterAgent.sinks.HDFS.hdfs.batchSize = 10000 //in github 1000
TwitterAgent.sinks.HDFS.hdfs.rollSize = 0
TwitterAgent.sinks.HDFS.hdfs.rollCount = 100000 //in github 10000

TwitterAgent.channels.MemChannel.type = memory
TwitterAgent.channels.MemChannel.capacity = 100000 //in github 10000
TwitterAgent.channels.MemChannel.transactionCapacity = 10000 //in github 100

或者 twitter 提供这条推文?而且不是hadoop问题?

UPD 1

这是我的flume配置:

# The configuration file needs to define the sources, 
# the channels and the sinks.
# Sources, channels and sinks are defined per agent, 
# in this case called 'TwitterAgent'

TwitterAgent.sources = Twitter
TwitterAgent.channels = MemChannel
TwitterAgent.sinks = HDFS

TwitterAgent.sources.Twitter.type = com.cloudera.flume.source.TwitterSource
 TwitterAgent.sources.Twitter.channels = MemChannel
 TwitterAgent.sources.Twitter.consumerKey = MyKey
 TwitterAgent.sources.Twitter.consumerSecret = MyKey
 TwitterAgent.sources.Twitter.accessToken = MyKey
 TwitterAgent.sources.Twitter.accessTokenSecret = MyKey
 TwitterAgent.sources.Twitter.keywords = hadoop, big-data , big data, analytics, bigdata, cloudera, data science, data scientiest, business intelligence, mapreduce, data warehouse, data warehousing, mahout, hbase, nosql, newsql, businessintelligence, cloudcomputing

 TwitterAgent.sinks.HDFS.channel = MemChannel
 TwitterAgent.sinks.HDFS.type = hdfs
 TwitterAgent.sinks.HDFS.hdfs.path = hdfs://rh-hadoop-master:8020/user/flume/tweets/%Y/%m/%d/%H/
 TwitterAgent.sinks.HDFS.hdfs.fileType = DataStream
 TwitterAgent.sinks.HDFS.hdfs.writeFormat = Text
 TwitterAgent.sinks.HDFS.hdfs.batchSize = 10000
 TwitterAgent.sinks.HDFS.hdfs.rollSize = 0
 TwitterAgent.sinks.HDFS.hdfs.rollCount = 100000

 TwitterAgent.channels.MemChannel.type = memory
 TwitterAgent.channels.MemChannel.capacity = 100000
 TwitterAgent.channels.MemChannel.transactionCapacity = 10000

这里是重复行的例子:

{"filter_level":"medium","retweeted":false,"in_reply_to_screen_name":null,"possibly_sensitive":false,"truncated":false,"lang":"en","in_reply_to_status_id_str":null,"id":539321584226680833,"in_reply_to_user_id_str":null,"timestamp_ms":"1417419260447","in_reply_to_status_id":null,"created_at":"Mon Dec 01 07:34:20 +0000 2014","favorite_count":0,"place":null,"coordinates":null,"text":"Testing Engineer, Hyderabad / Secunderabad, 2 - 5 Year Exp,Software Test Engineer , &amp;#x22;Big Data&amp;#x22;... http://t.co/DAK1ilWhM5","contributors":null,"geo":null,"entities":{"trends":[],"symbols":[],"urls":[{"expanded_url":"http://bit.ly/1ttBxPY","indices":[116,138],"display_url":"bit.ly/1ttBxPY","url":"http://t.co/DAK1ilWhM5"}],"hashtags":[{"text":"x22","indices":[89,93]},{"text":"x22","indices":[107,111]}],"user_mentions":[]},"source":"<a href=\"http://monsterindia.com\" rel=\"nofollow\">IT jobs, India<\/a>","favorited":false,"in_reply_to_user_id":null,"retweet_count":0,"id_str":"539321584226680833","user":{"location":"India","default_profile":false,"profile_background_tile":false,"statuses_count":63546,"lang":"en","profile_link_color":"0084B4","id":123537533,"following":null,"protected":false,"favourites_count":0,"profile_text_color":"333333","verified":false,"description":"Get latest job opportunities in Indian IT industry","contributors_enabled":false,"profile_sidebar_border_color":"C0DEED","name":"IT Jobs, India","profile_background_color":"C0DEED","created_at":"Tue Mar 16 11:48:44 +0000 2010","default_profile_image":false,"followers_count":1245,"profile_image_url_https":"https://pbs.twimg.com/profile_images/790482269/sm_it1_normal.jpg","geo_enabled":false,"profile_background_image_url":"http://pbs.twimg.com/profile_background_images/88067227/IT1.jpg","profile_background_image_url_https":"https://pbs.twimg.com/profile_background_images/88067227/IT1.jpg","follow_request_sent":null,"url":null,"utc_offset":null,"time_zone":null,"notifications":null,"profile_use_background_image":true,"friends_count":0,"profile_sidebar_fill_color":"DDEEF6","screen_name":"tech_career","id_str":"123537533","profile_image_url":"http://pbs.twimg.com/profile_images/790482269/sm_it1_normal.jpg","listed_count":43,"is_translator":false}}
{"filter_level":"medium","retweeted":false,"in_reply_to_screen_name":null,"possibly_sensitive":false,"truncated":false,"lang":"en","in_reply_to_status_id_str":null,"id":539321584226680833,"in_reply_to_user_id_str":null,"timestamp_ms":"1417419260447","in_reply_to_status_id":null,"created_at":"Mon Dec 01 07:34:20 +0000 2014","favorite_count":0,"place":null,"coordinates":null,"text":"Testing Engineer, Hyderabad / Secunderabad, 2 - 5 Year Exp,Software Test Engineer , &amp;#x22;Big Data&amp;#x22;... http://t.co/DAK1ilWhM5","contributors":null,"geo":null,"entities":{"trends":[],"symbols":[],"urls":[{"expanded_url":"http://bit.ly/1ttBxPY","indices":[116,138],"display_url":"bit.ly/1ttBxPY","url":"http://t.co/DAK1ilWhM5"}],"hashtags":[{"text":"x22","indices":[89,93]},{"text":"x22","indices":[107,111]}],"user_mentions":[]},"source":"<a href=\"http://monsterindia.com\" rel=\"nofollow\">IT jobs, India<\/a>","favorited":false,"in_reply_to_user_id":null,"retweet_count":0,"id_str":"539321584226680833","user":{"location":"India","default_profile":false,"profile_background_tile":false,"statuses_count":63546,"lang":"en","profile_link_color":"0084B4","id":123537533,"following":null,"protected":false,"favourites_count":0,"profile_text_color":"333333","verified":false,"description":"Get latest job opportunities in Indian IT industry","contributors_enabled":false,"profile_sidebar_border_color":"C0DEED","name":"IT Jobs, India","profile_background_color":"C0DEED","created_at":"Tue Mar 16 11:48:44 +0000 2010","default_profile_image":false,"followers_count":1245,"profile_image_url_https":"https://pbs.twimg.com/profile_images/790482269/sm_it1_normal.jpg","geo_enabled":false,"profile_background_image_url":"http://pbs.twimg.com/profile_background_images/88067227/IT1.jpg","profile_background_image_url_https":"https://pbs.twimg.com/profile_background_images/88067227/IT1.jpg","follow_request_sent":null,"url":null,"utc_offset":null,"time_zone":null,"notifications":null,"profile_use_background_image":true,"friends_count":0,"profile_sidebar_fill_color":"DDEEF6","screen_name":"tech_career","id_str":"123537533","profile_image_url":"http://pbs.twimg.com/profile_images/790482269/sm_it1_normal.jpg","listed_count":43,"is_translator":false}}

【问题讨论】:

    标签: hadoop twitter hadoop-streaming twitter-streaming-api


    【解决方案1】:

    Flume 不会向它要存储的数据添加任何类型的 id。 HDFS 也是如此,它在存储数据时不添加任何 id。它们只是一起工作,以便移动生成的数据并将其存储。

    如果您存储具有相同 id 的推文,那是因为您正在接收具有这些 id 的数据,或者您以错误的方式解释数据。

    话虽如此,也许您可​​以通过编辑来为您的问题添加一些示例。

    【讨论】:

    • 在存储的文件中(来自水槽)我可以看到多行相同的文本。
    • 能否将所有 Flume 配置添加到问题中,好吗?此外,您能否与我们分享您评论的那些复制行?
    • 完成。我添加了我的水槽配置和重复的行。
    • 我的赌注是同一条推文匹配了两个关键字,但我只能在您打印的推文中找到“大数据”...您可以通过配置唯一关键字进行测试吗?只是为了确认即使在这种情况下,您也在 HDFS 中复制推文。
    • 我从这个流媒体的一开始就制作了一些带有多个主题标签的测试推文。但每次我在 Hive 表中只收到一条推文。
    猜你喜欢
    • 2017-06-09
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2014-03-21
    • 2021-12-04
    • 2023-03-08
    • 2018-02-11
    相关资源
    最近更新 更多