【问题标题】:How to get HTML content of webpage from ASP.NET如何从 ASP.NET 获取网页的 HTML 内容
【发布时间】:2014-09-23 19:02:16
【问题描述】:

我想从动态网页中抓取一些内容(好像是用 MVC 开发的)。

数据抓取逻辑是通过 HTML 敏捷性完成的,但现在的问题是, 从浏览器请求 URL 时返回的 HTML 与来自 ASP.NET Web 请求的 URL 的 Web 响应不同。

主要是浏览器响应有我需要的动态数据(根据查询字符串中传递的值呈现),但WebResponse结果不同。

能否请您帮我获取动态网页视图WebRequest的实际内容。

下面是我以前看过的代码:

WebRequest request = WebRequest.Create(sURL);
request.Method = "Get";
//Get the response
WebResponse response = request.GetResponse();
//Read the stream from the response
StreamReader reader = new StreamReader(response.GetResponseStream(), System.Text.Encoding.UTF8);

【问题讨论】:

  • “WebResponse 结果”有何不同?
  • 我从 Web URL 的 View Source 获得的 HTML 内容与 ASP.Net Web 响应的 HTML 内容不同。例如:如果 URL 用于特定邮政编码中的餐馆,则浏览器请求的 HTML 具有该区域的餐馆列表,但 Web 响应 HTML 在该 DIV 中没有餐馆。
  • 餐厅 1
  • 餐厅 2
-- 浏览器响应。
  • -- ASP.Net 网络响应

    标签: c# html asp.net asp.net-mvc httpwebrequest


    【解决方案1】:

    使用HttpWebRequest获取任何网页的内容...

    // We will store the html response of the request here
    string siteContent = string.Empty;
    
    // The url you want to grab
    string url = "http://google.com";
    
    // Here we're creating our request, we haven't actually sent the request to the site yet...
    // we're simply building our HTTP request to shoot off to google...
    HttpWebRequest request = (HttpWebRequest)WebRequest.Create(url);
    request.AutomaticDecompression = DecompressionMethods.GZip;
    
    // Right now... this is what our HTTP Request has been built in to...
    /*
        GET http://google.com/ HTTP/1.1
        Host: google.com
        Accept-Encoding: gzip
        Connection: Keep-Alive
    */
    
    
    // Wrap everything that can be disposed in using blocks... 
    // They dispose of objects and prevent them from lying around in memory...
    using(HttpWebResponse response = (HttpWebResponse)request.GetResponse())  // Go query google
    using(Stream responseStream = response.GetResponseStream())               // Load the response stream
    using(StreamReader streamReader = new StreamReader(responseStream))       // Load the stream reader to read the response
    {
        siteContent = streamReader.ReadToEnd(); // Read the entire response and store it in the siteContent variable
    }
    
    // magic...
    Console.WriteLine (siteContent);
    

    【讨论】:

    • 如果你要否决它...写一个评论至少解释原因以便我可以改进它,而其他人不会觉得这个解决方案不起作用
    • 谢谢你,但我只需要知道为什么来自浏览器的 URL 请求返​​回正确的结果 (HTML),而 ASP.Net Web 请求不能。这是因为请求的站点是在 ASP.Net MVC 中开发的吗?
    • 它将能够...可能还有其他事情发生,但无论是什么...我可以向您保证,如果您可以在网络浏览器中访问它...您可以使用 HttpWebRequest 类访问它,因为它们都使用 HTTP。问题可能是通过 JavaScript 进行重定向...或者可能在您请求的页面加载后触发 AJAX 请求...下载 Fiddler...并查看访问页面后发生的 HTTP 请求在浏览器中
    • 这非常有效。它通过客户端执行此请求。所以无需在客户端使用 jquery 等进行尝试。
    猜你喜欢
    相关资源
    最近更新 更多
    热门标签