使用jsoup下载大的pdf [英] Download a large pdf with jsoup

查看:290
本文介绍了使用jsoup下载大的pdf的处理方法,对大家解决问题具有一定的参考价值,需要的朋友们下面随着小编来一起学习吧!

问题描述

我想用jsoup下载大的pdf文件.我尝试更改超时和maxBodySize,但我可以下载的最大文件约为11MB.我认为是否有任何办法可以做诸如缓冲之类的事情.下面是我的代码.

I would like to download a large pdf file with jsoup. I have try to change timeout and maxBodySize but the largest file I could download was about 11MB. I think if there is any way to do something like buffering. Below is my code.

public class Download extends Activity {

static public String nextPage;
static public Response file;
static public Connection.Response res;

@Override
protected void onCreate(Bundle savedInstanceState) {
    // TODO Auto-generated method stub
    super.onCreate(savedInstanceState);
    Bundle b = new Bundle();
    b = getIntent().getExtras();
    nextPage = b.getString("key");
    new Login().execute();
    finish();
}

private class Login extends AsyncTask<Void, Void, Void> {

    @Override
    protected void onPreExecute() {
        super.onPreExecute();
    }

    @Override
    protected Void doInBackground(Void... params) {
        try {
            res = Jsoup.connect("http://www.eclass.teikal.gr/eclass2/")
                    .ignoreContentType(true).userAgent("Mozilla/5.0")
                    .execute();

            SharedPreferences pref = getSharedPreferences(
                    MainActivity.PREFS_NAME, MODE_PRIVATE);
            String username1 = pref.getString(MainActivity.PREF_USERNAME,
                    null);
            String password1 = pref.getString(MainActivity.PREF_PASSWORD,
                    null);
            file = (Response) Jsoup
                    .connect("http://www.eclass.teikal.gr/eclass2/")
                    .ignoreContentType(true).userAgent("Mozilla/5.0")
                    .maxBodySize(1024*1024*10*2)
                    .timeout(70000*10)
                    .cookies(res.cookies()).data("uname", username1)
                    .data("pass", password1).data("next", nextPage)
                    .data("submit", "").method(Method.POST).execute();

        } catch (IOException e) {
            e.printStackTrace();
        }
        return null;

    }

    @Override
    protected void onPostExecute(Void result) {

        String PATH = Environment.getExternalStorageDirectory()
                + "/download/";
        String name = "eclassTest.pdf";
        FileOutputStream out;
        try {

            int len = file.bodyAsBytes().length;
            out = new FileOutputStream(new File(PATH + name));
            out.write(file.bodyAsBytes(),0,len);
            out.close();
        } catch (FileNotFoundException e) {
            e.printStackTrace();
        } catch (IOException e) {
            e.printStackTrace();
        }

    }
  }
}

我希望有人能帮助我!

推荐答案

我认为,最好通过HTTPConnection下载任何二进制文件:

I think, it's better to download any binary file via HTTPConnection:

    InputStream input = null;
    OutputStream output = null;
    HttpURLConnection connection = null;
    try {
        URL url = new URL("http://example.com/file.pdf");
        connection = (HttpURLConnection) url.openConnection();
        connection.connect();

        // expect HTTP 200 OK, so we don't mistakenly save error report
        // instead of the file
        if (connection.getResponseCode() != HttpURLConnection.HTTP_OK) {
            return "Server returned HTTP " + connection.getResponseCode()
                    + " " + connection.getResponseMessage();
        }

        // this will be useful to display download percentage
        // might be -1: server did not report the length
        int fileLength = connection.getContentLength();

        // download the file
        input = connection.getInputStream();
        output = new FileOutputStream("/sdcard/file_name.extension");

        byte data[] = new byte[4096];
        int count;
        while ((count = input.read(data)) != -1) {
            output.write(data, 0, count);
        }
    } catch (Exception e) {
        return e.toString();
    } finally {
        try {
            if (output != null)
                output.close();
            if (input != null)
                input.close();
        } catch (IOException ignored) {
        }

        if (connection != null)
            connection.disconnect();
    }

Jsoup用于解析和加载HTML页面,而不是二进制文件.

Jsoup is for parsing and loading HTML pages, not binary files.

这篇关于使用jsoup下载大的pdf的文章就介绍到这了,希望我们推荐的答案对大家有所帮助,也希望大家多多支持IT屋!

查看全文
登录 关闭
扫码关注1秒登录
发送“验证码”获取 | 15天全站免登陆