草庐IT

xml - 如何将xml文件转换为R中的数据框

coder 2024-06-26 原文

我有这个 xml 文件,我想将它转换为数据框:

数据.xml

 <?xml version="1.0" encoding="UTF-8" standalone="yes" ?> 
- <graph_data xmlns:ns2="http://www.w3.org/2005/Atom">
  <graph_property name="calculation_method" value="Geo Mean" /> 
  <graph_property name="graph_type" value="TIME" /> 
- <measurement id="521406">
  <alias>site4</alias> 
- <bucket_data>
- <bucket id="1" name="2013-MAY-14 07:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="21" /> 
  <perf_data unit="seconds" value="3.102" /> 
  </bucket>
- <bucket id="2" name="2013-MAY-14 08:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="13" /> 
  <perf_data unit="seconds" value="3.052" /> 
  </bucket>
- <bucket id="3" name="2013-MAY-14 09:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="3.387" /> 
  </bucket>
- <bucket id="4" name="2013-MAY-14 10:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="3.338" /> 
  </bucket>
- <bucket id="5" name="2013-MAY-14 11:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="10" /> 
  <perf_data unit="seconds" value="2.149" /> 
  </bucket>
- <bucket id="6" name="2013-MAY-14 12:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="13" /> 
  <perf_data unit="seconds" value="3.202" /> 
  </bucket>
- <bucket id="7" name="2013-MAY-14 01:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="18" /> 
  <perf_data unit="seconds" value="2.883" /> 
  </bucket>
- <bucket id="8" name="2013-MAY-14 02:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="11" /> 
  <perf_data unit="seconds" value="2.582" /> 
  </bucket>
- <bucket id="9" name="2013-MAY-14 03:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="2.769" /> 
  </bucket>
- <bucket id="10" name="2013-MAY-14 04:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="12" /> 
  <perf_data unit="seconds" value="2.669" /> 
  </bucket>
- <bucket id="11" name="2013-MAY-14 05:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="17" /> 
  <perf_data unit="seconds" value="2.830" /> 
  </bucket>
- <bucket id="12" name="2013-MAY-14 06:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="12" /> 
  <perf_data unit="seconds" value="2.591" /> 
  </bucket>
- <bucket id="13" name="2013-MAY-14 07:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="17" /> 
  <perf_data unit="seconds" value="3.213" /> 
  </bucket>
- <bucket id="14" name="2013-MAY-14 08:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="10" /> 
  <perf_data unit="seconds" value="2.653" /> 
  </bucket>
- <bucket id="15" name="2013-MAY-14 09:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="2.935" /> 
  </bucket>
- <bucket id="16" name="2013-MAY-14 10:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="18" /> 
  <perf_data unit="seconds" value="2.495" /> 
  </bucket>
- <bucket id="17" name="2013-MAY-14 11:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="10" /> 
  <perf_data unit="seconds" value="3.169" /> 
  </bucket>
- <bucket id="18" name="2013-MAY-15 12:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="16" /> 
  <perf_data unit="seconds" value="2.789" /> 
  </bucket>
- <bucket id="19" name="2013-MAY-15 01:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="16" /> 
  <perf_data unit="seconds" value="3.245" /> 
  </bucket>
- <bucket id="20" name="2013-MAY-15 02:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="17" /> 
  <perf_data unit="seconds" value="3.281" /> 
  </bucket>
- <bucket id="21" name="2013-MAY-15 03:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="12" /> 
  <perf_data unit="seconds" value="3.773" /> 
  </bucket>
- <bucket id="22" name="2013-MAY-15 04:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="21" /> 
  <perf_data unit="seconds" value="2.648" /> 
  </bucket>
- <bucket id="23" name="2013-MAY-15 05:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="11" /> 
  <perf_data unit="seconds" value="3.291" /> 
  </bucket>
- <bucket id="24" name="2013-MAY-15 06:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="12" /> 
  <perf_data unit="seconds" value="3.084" /> 
  </bucket>
  </bucket_data>
- <graph_option>
  <data_cell name="perfwarning" unit="seconds" value="-" /> 
  <data_cell name="perfcritical" unit="seconds" value="-" /> 
  <data_cell name="availwarning" unit="percent" value="-" /> 
  <data_cell name="availcritical" unit="percent" value="-" /> 
  <data_cell name="bucketsize" unit="seconds" value="3600" /> 
  <data_cell name="rows" unit="#" value="24" /> 
  <data_cell name="pagecomponent" unit="seconds" value="Total Time" /> 
  <data_cell name="avg_perf" unit="seconds" value="2.949" /> 
  <data_cell name="avg_avail" unit="percent" value="100.00" /> 
  <data_cell name="total_datapoint_count" unit="#" value="347" /> 
  <data_cell /> 
  </graph_option>
  </measurement>
- <measurement id="521406">
  <alias>Site3</alias> 
- <bucket_data>
- <bucket id="1" name="2013-MAY-14 07:00 AM">
  <avail_data unit="percent" value="85.71" /> 
  <data_count unit="#" value="18" /> 
  <perf_data unit="seconds" value="6.503" /> 
  </bucket>
- <bucket id="2" name="2013-MAY-14 08:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="13" /> 
  <perf_data unit="seconds" value="6.330" /> 
  </bucket>
- <bucket id="3" name="2013-MAY-14 09:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="7.242" /> 
  </bucket>
- <bucket id="4" name="2013-MAY-14 10:00 AM">
  <avail_data unit="percent" value="93.33" /> 
  <data_count unit="#" value="14" /> 
  <perf_data unit="seconds" value="7.083" /> 
  </bucket>
- <bucket id="5" name="2013-MAY-14 11:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="10" /> 
  <perf_data unit="seconds" value="7.087" /> 
  </bucket>
- <bucket id="6" name="2013-MAY-14 12:00 PM">
  <avail_data unit="percent" value="76.92" /> 
  <data_count unit="#" value="10" /> 
  <perf_data unit="seconds" value="6.197" /> 
  </bucket>
- <bucket id="7" name="2013-MAY-14 01:00 PM">
  <avail_data unit="percent" value="83.33" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="6.772" /> 
  </bucket>
- <bucket id="8" name="2013-MAY-14 02:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="11" /> 
  <perf_data unit="seconds" value="5.832" /> 
  </bucket>
- <bucket id="9" name="2013-MAY-14 03:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="8.513" /> 
  </bucket>
- <bucket id="10" name="2013-MAY-14 04:00 PM">
  <avail_data unit="percent" value="91.67" /> 
  <data_count unit="#" value="11" /> 
  <perf_data unit="seconds" value="7.190" /> 
  </bucket>
- <bucket id="11" name="2013-MAY-14 05:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="17" /> 
  <perf_data unit="seconds" value="6.373" /> 
  </bucket>
- <bucket id="12" name="2013-MAY-14 06:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="12" /> 
  <perf_data unit="seconds" value="8.440" /> 
  </bucket>
- <bucket id="13" name="2013-MAY-14 07:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="17" /> 
  <perf_data unit="seconds" value="6.318" /> 
  </bucket>
- <bucket id="14" name="2013-MAY-14 08:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="10" /> 
  <perf_data unit="seconds" value="6.374" /> 
  </bucket>
- <bucket id="15" name="2013-MAY-14 09:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="6.773" /> 
  </bucket>
- <bucket id="16" name="2013-MAY-14 10:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="18" /> 
  <perf_data unit="seconds" value="6.274" /> 
  </bucket>
- <bucket id="17" name="2013-MAY-14 11:00 PM">
  <avail_data unit="percent" value="90.00" /> 
  <data_count unit="#" value="9" /> 
  <perf_data unit="seconds" value="5.881" /> 
  </bucket>
- <bucket id="18" name="2013-MAY-15 12:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="16" /> 
  <perf_data unit="seconds" value="5.630" /> 
  </bucket>
- <bucket id="19" name="2013-MAY-15 01:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="16" /> 
  <perf_data unit="seconds" value="6.585" /> 
  </bucket>
- <bucket id="20" name="2013-MAY-15 02:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="17" /> 
  <perf_data unit="seconds" value="7.394" /> 
  </bucket>
- <bucket id="21" name="2013-MAY-15 03:00 AM">
  <avail_data unit="percent" value="91.67" /> 
  <data_count unit="#" value="11" /> 
  <perf_data unit="seconds" value="6.427" /> 
  </bucket>
- <bucket id="22" name="2013-MAY-15 04:00 AM">
  <avail_data unit="percent" value="95.24" /> 
  <data_count unit="#" value="20" /> 
  <perf_data unit="seconds" value="7.140" /> 
  </bucket>
- <bucket id="23" name="2013-MAY-15 05:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="11" /> 
  <perf_data unit="seconds" value="7.152" /> 
  </bucket>
- <bucket id="24" name="2013-MAY-15 06:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="12" /> 
  <perf_data unit="seconds" value="6.474" /> 
  </bucket>
  </bucket_data>
- <graph_option>
  <data_cell name="perfwarning" unit="seconds" value="-" /> 
  <data_cell name="perfcritical" unit="seconds" value="-" /> 
  <data_cell name="availwarning" unit="percent" value="-" /> 
  <data_cell name="availcritical" unit="percent" value="-" /> 
  <data_cell name="bucketsize" unit="seconds" value="3600" /> 
  <data_cell name="rows" unit="#" value="24" /> 
  <data_cell name="pagecomponent" unit="seconds" value="Total Time" /> 
  <data_cell name="avg_perf" unit="seconds" value="6.729" /> 
  <data_cell name="avg_avail" unit="percent" value="95.97" /> 
  <data_cell name="total_datapoint_count" unit="#" value="347" /> 
  <data_cell /> 
  </graph_option>
  </measurement>
- <measurement id="521406">
  <alias>Site2</alias> 
- <bucket_data>
- <bucket id="1" name="2013-MAY-14 07:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="18" /> 
  <perf_data unit="seconds" value="2.247" /> 
  </bucket>
- <bucket id="2" name="2013-MAY-14 08:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="13" /> 
  <perf_data unit="seconds" value="2.382" /> 
  </bucket>
- <bucket id="3" name="2013-MAY-14 09:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="2.232" /> 
  </bucket>
- <bucket id="4" name="2013-MAY-14 10:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="14" /> 
  <perf_data unit="seconds" value="2.223" /> 
  </bucket>
- <bucket id="5" name="2013-MAY-14 11:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="10" /> 
  <perf_data unit="seconds" value="2.265" /> 
  </bucket>
- <bucket id="6" name="2013-MAY-14 12:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="10" /> 
  <perf_data unit="seconds" value="2.130" /> 
  </bucket>
- <bucket id="7" name="2013-MAY-14 01:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="2.153" /> 
  </bucket>
- <bucket id="8" name="2013-MAY-14 02:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="11" /> 
  <perf_data unit="seconds" value="2.005" /> 
  </bucket>
- <bucket id="9" name="2013-MAY-14 03:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="2.322" /> 
  </bucket>
- <bucket id="10" name="2013-MAY-14 04:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="11" /> 
  <perf_data unit="seconds" value="1.918" /> 
  </bucket>
- <bucket id="11" name="2013-MAY-14 05:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="17" /> 
  <perf_data unit="seconds" value="1.992" /> 
  </bucket>
- <bucket id="12" name="2013-MAY-14 06:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="12" /> 
  <perf_data unit="seconds" value="2.423" /> 
  </bucket>
- <bucket id="13" name="2013-MAY-14 07:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="17" /> 
  <perf_data unit="seconds" value="2.327" /> 
  </bucket>
- <bucket id="14" name="2013-MAY-14 08:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="10" /> 
  <perf_data unit="seconds" value="2.605" /> 
  </bucket>
- <bucket id="15" name="2013-MAY-14 09:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="2.533" /> 
  </bucket>
- <bucket id="16" name="2013-MAY-14 10:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="18" /> 
  <perf_data unit="seconds" value="2.077" /> 
  </bucket>
- <bucket id="17" name="2013-MAY-14 11:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="9" /> 
  <perf_data unit="seconds" value="2.356" /> 
  </bucket>
- <bucket id="18" name="2013-MAY-15 12:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="16" /> 
  <perf_data unit="seconds" value="2.506" /> 
  </bucket>
- <bucket id="19" name="2013-MAY-15 01:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="16" /> 
  <perf_data unit="seconds" value="2.422" /> 
  </bucket>
- <bucket id="20" name="2013-MAY-15 02:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="17" /> 
  <perf_data unit="seconds" value="2.220" /> 
  </bucket>
- <bucket id="21" name="2013-MAY-15 03:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="11" /> 
  <perf_data unit="seconds" value="2.669" /> 
  </bucket>
- <bucket id="22" name="2013-MAY-15 04:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="20" /> 
  <perf_data unit="seconds" value="2.274" /> 
  </bucket>
- <bucket id="23" name="2013-MAY-15 05:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="11" /> 
  <perf_data unit="seconds" value="2.277" /> 
  </bucket>
- <bucket id="24" name="2013-MAY-15 06:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="12" /> 
  <perf_data unit="seconds" value="2.180" /> 
  </bucket>
  </bucket_data>
- <graph_option>
  <data_cell name="perfwarning" unit="seconds" value="-" /> 
  <data_cell name="perfcritical" unit="seconds" value="-" /> 
  <data_cell name="availwarning" unit="percent" value="-" /> 
  <data_cell name="availcritical" unit="percent" value="-" /> 
  <data_cell name="bucketsize" unit="seconds" value="3600" /> 
  <data_cell name="rows" unit="#" value="24" /> 
  <data_cell name="pagecomponent" unit="seconds" value="Total Time" /> 
  <data_cell name="avg_perf" unit="seconds" value="2.269" /> 
  <data_cell name="avg_avail" unit="percent" value="100.00" /> 
  <data_cell name="total_datapoint_count" unit="#" value="333" /> 
  <data_cell /> 
  </graph_option>
  </measurement>
- <measurement id="521406">
  <alias>Site1</alias> 
- <bucket_data>
- <bucket id="1" name="2013-MAY-14 07:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="18" /> 
  <perf_data unit="seconds" value="1.431" /> 
  </bucket>
- <bucket id="2" name="2013-MAY-14 08:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="13" /> 
  <perf_data unit="seconds" value="1.559" /> 
  </bucket>
- <bucket id="3" name="2013-MAY-14 09:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="1.378" /> 
  </bucket>
- <bucket id="4" name="2013-MAY-14 10:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="14" /> 
  <perf_data unit="seconds" value="1.307" /> 
  </bucket>
- <bucket id="5" name="2013-MAY-14 11:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="10" /> 
  <perf_data unit="seconds" value="1.458" /> 
  </bucket>
- <bucket id="6" name="2013-MAY-14 12:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="10" /> 
  <perf_data unit="seconds" value="1.345" /> 
  </bucket>
- <bucket id="7" name="2013-MAY-14 01:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="1.317" /> 
  </bucket>
- <bucket id="8" name="2013-MAY-14 02:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="11" /> 
  <perf_data unit="seconds" value="1.465" /> 
  </bucket>
- <bucket id="9" name="2013-MAY-14 03:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="1.398" /> 
  </bucket>
- <bucket id="10" name="2013-MAY-14 04:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="11" /> 
  <perf_data unit="seconds" value="1.509" /> 
  </bucket>
- <bucket id="11" name="2013-MAY-14 05:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="17" /> 
  <perf_data unit="seconds" value="1.284" /> 
  </bucket>
- <bucket id="12" name="2013-MAY-14 06:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="12" /> 
  <perf_data unit="seconds" value="1.759" /> 
  </bucket>
- <bucket id="13" name="2013-MAY-14 07:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="17" /> 
  <perf_data unit="seconds" value="1.434" /> 
  </bucket>
- <bucket id="14" name="2013-MAY-14 08:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="10" /> 
  <perf_data unit="seconds" value="1.402" /> 
  </bucket>
- <bucket id="15" name="2013-MAY-14 09:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="1.452" /> 
  </bucket>
- <bucket id="16" name="2013-MAY-14 10:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="18" /> 
  <perf_data unit="seconds" value="1.216" /> 
  </bucket>
- <bucket id="17" name="2013-MAY-14 11:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="9" /> 
  <perf_data unit="seconds" value="1.381" /> 
  </bucket>
- <bucket id="18" name="2013-MAY-15 12:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="16" /> 
  <perf_data unit="seconds" value="1.236" /> 
  </bucket>
- <bucket id="19" name="2013-MAY-15 01:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="16" /> 
  <perf_data unit="seconds" value="1.327" /> 
  </bucket>
- <bucket id="20" name="2013-MAY-15 02:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="17" /> 
  <perf_data unit="seconds" value="1.465" /> 
  </bucket>
- <bucket id="21" name="2013-MAY-15 03:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="11" /> 
  <perf_data unit="seconds" value="1.529" /> 
  </bucket>
- <bucket id="22" name="2013-MAY-15 04:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="20" /> 
  <perf_data unit="seconds" value="1.354" /> 
  </bucket>
- <bucket id="23" name="2013-MAY-15 05:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="11" /> 
  <perf_data unit="seconds" value="1.372" /> 
  </bucket>
- <bucket id="24" name="2013-MAY-15 06:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="12" /> 
  <perf_data unit="seconds" value="1.219" /> 
  </bucket>
  </bucket_data>
- <graph_option>
  <data_cell name="perfwarning" unit="seconds" value="-" /> 
  <data_cell name="perfcritical" unit="seconds" value="-" /> 
  <data_cell name="availwarning" unit="percent" value="-" /> 
  <data_cell name="availcritical" unit="percent" value="-" /> 
  <data_cell name="bucketsize" unit="seconds" value="3600" /> 
  <data_cell name="rows" unit="#" value="24" /> 
  <data_cell name="pagecomponent" unit="seconds" value="Total Time" /> 
  <data_cell name="avg_perf" unit="seconds" value="1.387" /> 
  <data_cell name="avg_avail" unit="percent" value="100.00" /> 
  <data_cell name="total_datapoint_count" unit="#" value="333" /> 
  <data_cell /> 
  </graph_option>
  </measurement>
  <ns2:link href="www.example.com" rel="slotmetadata" type="application/xml" /> 
  </graph_data>

我的数据框需要如下所示:

alias  bucket_name avail_data perf_data

我试过这个:

doc1 = xmlParse("data.xml")
df<-xmlToDataFrame(nodes = getNodeSet(doc1, "//alias"))

我只在一列数据框中获取别名。任何想法我在这里还缺少什么?

有文件

最佳答案

看起来您的 XML 有一些问题。我只能通过删除以下内容来让它工作:

Line 58: <measurement id="521406">
Line 107: <measurement id="521406">

所以:

xml_file <- '<?xml version="1.0" encoding="UTF-8" standalone="yes" ?> 
  <graph_data xmlns:ns2="http://www.w3.org/2005/Atom">
  <graph_property name="calculation_method" value="Geo Mean" /> 
  <graph_property name="graph_type" value="TIME" /> <measurement id="521406">
  <alias>example1.com</alias> 
  <bucket_data>
  <bucket id="1" name="2013-MAY-14 07:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="21" /> 
  <perf_data unit="seconds" value="3.102" /> 
  </bucket>
  <bucket id="2" name="2013-MAY-14 08:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="13" /> 
  <perf_data unit="seconds" value="3.052" /> 
  </bucket>
  <bucket id="3" name="2013-MAY-14 09:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="3.387" /> 
  </bucket>
  <bucket id="4" name="2013-MAY-14 10:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="3.338" /> 
  </bucket>
  <bucket id="5" name="2013-MAY-14 11:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="10" /> 
  <perf_data unit="seconds" value="2.149" /> 
  </bucket>
  <bucket id="6" name="2013-MAY-14 12:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="13" /> 
  <perf_data unit="seconds" value="3.202" /> 
  </bucket>
  <bucket id="7" name="2013-MAY-14 01:00 PM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="18" /> 
  <perf_data unit="seconds" value="2.883" /> 
  </bucket>
  </bucket_data>
  <graph_option>
  <data_cell name="perfwarning" unit="seconds" value="-" /> 
  <data_cell name="perfcritical" unit="seconds" value="-" /> 
  <data_cell name="availwarning" unit="percent" value="-" /> 
  <data_cell name="availcritical" unit="percent" value="-" /> 
  <data_cell name="bucketsize" unit="seconds" value="3600" /> 
  <data_cell name="rows" unit="#" value="24" /> 
  <data_cell name="pagecomponent" unit="seconds" value="Total Time" /> 
  <data_cell name="avg_perf" unit="seconds" value="2.949" /> 
  <data_cell name="avg_avail" unit="percent" value="100.00" /> 
  <data_cell name="total_datapoint_count" unit="#" value="347" /> 
  <data_cell /> 
  </graph_option>
  </measurement>
  <measurement id="521406">
  <alias>example2.com</alias> 
  <bucket_data>
  <bucket id="1" name="2013-MAY-14 07:00 AM">
  <avail_data unit="percent" value="85.71" /> 
  <data_count unit="#" value="18" /> 
  <perf_data unit="seconds" value="6.503" /> 
  </bucket>
  <bucket id="2" name="2013-MAY-14 08:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="13" /> 
  <perf_data unit="seconds" value="6.330" /> 
  </bucket>
  <bucket id="3" name="2013-MAY-14 09:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="15" /> 
  <perf_data unit="seconds" value="7.242" /> 
  </bucket>
  <bucket id="4" name="2013-MAY-14 10:00 AM">
  <avail_data unit="percent" value="93.33" /> 
  <data_count unit="#" value="14" /> 
  <perf_data unit="seconds" value="7.083" /> 
  </bucket>
  <bucket id="5" name="2013-MAY-14 11:00 AM">
  <avail_data unit="percent" value="100.00" /> 
  <data_count unit="#" value="10" /> 
  <perf_data unit="seconds" value="7.087" /> 
  </bucket>
  <bucket id="6" name="2013-MAY-14 12:00 PM">
  <avail_data unit="percent" value="76.92" /> 
  <data_count unit="#" value="10" /> 
  <perf_data unit="seconds" value="6.197" /> 
  </bucket>
  </bucket_data>
  <graph_option>
  <data_cell name="perfwarning" unit="seconds" value="-" /> 
  <data_cell name="perfcritical" unit="seconds" value="-" /> 
  <data_cell name="availwarning" unit="percent" value="-" /> 
  <data_cell name="availcritical" unit="percent" value="-" /> 
  <data_cell name="bucketsize" unit="seconds" value="3600" /> 
  <data_cell name="rows" unit="#" value="24" /> 
  <data_cell name="pagecomponent" unit="seconds" value="Total Time" /> 
  <data_cell name="avg_perf" unit="seconds" value="6.729" /> 
  <data_cell name="avg_avail" unit="percent" value="95.97" /> 
  <data_cell name="total_datapoint_count" unit="#" value="347" /> 
  <data_cell /> 
  </graph_option>
  </measurement>
  </graph_data>'

xml_file <- xmlParse(xml_file)    # Parse the XML
xml_file <- xmlToList(xml_file)   # Convert the XML to a list

我将其转换为列表而不是数据框,因为 XML 似乎不遵循可轻松转换为行和列的结构。之后,根据你的问题,我只提取了“测量”节点的“别名”或“bucket_data”部分包含的信息:

xml_file <- xml_file[names(xml_file) == "measurement"]
xml_file <- lapply(xml_file, function(x) x[grep("alias|bucket", names(x))])

然后我遍历每个测量节点,搁置别名信息,将桶列表变成一个命名向量,然后将别名和桶绑定(bind)在一起成为列。最后,我将测量节点绑定(bind)到行中并将整个事物转换为数据框。

xml_file <- lapply(xml_file, function(x) {
  alias <- x$alias
  buckets <- t(sapply(x$bucket_data, unlist))
  cbind("alias" = alias, buckets)
})

xml_file <- do.call("rbind", xml_file)

xml_file <- data.frame(xml_file, stringsAsFactors = FALSE)
Warning message:
In data.row.names(row.names, rowsi, i) :
  some row.names duplicated: 2,3,4,5,6,7,8,9,10,11,12,13 --> row.names NOT used

str(xml_file)
'data.frame':   13 obs. of  9 variables:
 $ alias           : chr  "example1.com" "example1.com" "example1.com" "example1.com" ...
 $ avail_data.unit : chr  "percent" "percent" "percent" "percent" ...
 $ avail_data.value: chr  "100.00" "100.00" "100.00" "100.00" ...
 $ data_count.unit : chr  "#" "#" "#" "#" ...
 $ data_count.value: chr  "21" "13" "15" "15" ...
 $ perf_data.unit  : chr  "seconds" "seconds" "seconds" "seconds" ...
 $ perf_data.value : chr  "3.102" "3.052" "3.387" "3.338" ...
 $ .attrs.id       : chr  "1" "2" "3" "4" ...
 $ .attrs.name     : chr  "2013-MAY-14 07:00 AM" "2013-MAY-14 08:00 AM" "2013-MAY-14 09:00 AM" "2013-MAY-14 10:00 AM" ...

您仍然需要清理列名称并将列转换为适当的类,但它会将您的数据放入数据框中。

关于xml - 如何将xml文件转换为R中的数据框,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/16574764/

有关xml - 如何将xml文件转换为R中的数据框的更多相关文章

  1. ruby - 如何使用 Nokogiri 的 xpath 和 at_xpath 方法 - 2

    我正在学习如何使用Nokogiri,根据这段代码我遇到了一些问题:require'rubygems'require'mechanize'post_agent=WWW::Mechanize.newpost_page=post_agent.get('http://www.vbulletin.org/forum/showthread.php?t=230708')puts"\nabsolutepathwithtbodygivesnil"putspost_page.parser.xpath('/html/body/div/div/div/div/div/table/tbody/tr/td/div

  2. ruby - 如何从 ruby​​ 中的字符串运行任意对象方法? - 2

    总的来说,我对ruby​​还比较陌生,我正在为我正在创建的对象编写一些rspec测试用例。许多测试用例都非常基础,我只是想确保正确填充和返回值。我想知道是否有办法使用循环结构来执行此操作。不必为我要测试的每个方法都设置一个assertEquals。例如:describeitem,"TestingtheItem"doit"willhaveanullvaluetostart"doitem=Item.new#HereIcoulddotheitem.name.shouldbe_nil#thenIcoulddoitem.category.shouldbe_nilendend但我想要一些方法来使用

  3. ruby - 使用 RubyZip 生成 ZIP 文件时设置压缩级别 - 2

    我有一个Ruby程序,它使用rubyzip压缩XML文件的目录树。gem。我的问题是文件开始变得很重,我想提高压缩级别,因为压缩时间不是问题。我在rubyzipdocumentation中找不到一种为创建的ZIP文件指定压缩级别的方法。有人知道如何更改此设置吗?是否有另一个允许指定压缩级别的Ruby库? 最佳答案 这是我通过查看ruby​​zip内部创建的代码。level=Zlib::BEST_COMPRESSIONZip::ZipOutputStream.open(zip_file)do|zip|Dir.glob("**/*")d

  4. ruby - 其他文件中的 Rake 任务 - 2

    我试图在一个项目中使用rake,如果我把所有东西都放到Rakefile中,它会很大并且很难读取/找到东西,所以我试着将每个命名空间放在lib/rake中它自己的文件中,我添加了这个到我的rake文件的顶部:Dir['#{File.dirname(__FILE__)}/lib/rake/*.rake'].map{|f|requiref}它加载文件没问题,但没有任务。我现在只有一个.rake文件作为测试,名为“servers.rake”,它看起来像这样:namespace:serverdotask:testdoputs"test"endend所以当我运行rakeserver:testid时

  5. ruby-on-rails - 在 Rails 中将文件大小字符串转换为等效千字节 - 2

    我的目标是转换表单输入,例如“100兆字节”或“1GB”,并将其转换为我可以存储在数据库中的文件大小(以千字节为单位)。目前,我有这个:defquota_convert@regex=/([0-9]+)(.*)s/@sizes=%w{kilobytemegabytegigabyte}m=self.quota.match(@regex)if@sizes.include?m[2]eval("self.quota=#{m[1]}.#{m[2]}")endend这有效,但前提是输入是倍数(“gigabytes”,而不是“gigabyte”)并且由于使用了eval看起来疯狂不安全。所以,功能正常,

  6. ruby-on-rails - Ruby net/ldap 模块中的内存泄漏 - 2

    作为我的Rails应用程序的一部分,我编写了一个小导入程序,它从我们的LDAP系统中吸取数据并将其塞入一个用户表中。不幸的是,与LDAP相关的代码在遍历我们的32K用户时泄漏了大量内存,我一直无法弄清楚如何解决这个问题。这个问题似乎在某种程度上与LDAP库有关,因为当我删除对LDAP内容的调用时,内存使用情况会很好地稳定下来。此外,不断增加的对象是Net::BER::BerIdentifiedString和Net::BER::BerIdentifiedArray,它们都是LDAP库的一部分。当我运行导入时,内存使用量最终达到超过1GB的峰值。如果问题存在,我需要找到一些方法来更正我的代

  7. python - 如何使用 Ruby 或 Python 创建一系列高音调和低音调的蜂鸣声? - 2

    关闭。这个问题是opinion-based.它目前不接受答案。想要改进这个问题?更新问题,以便editingthispost可以用事实和引用来回答它.关闭4年前。Improvethisquestion我想在固定时间创建一系列低音和高音调的哔哔声。例如:在150毫秒时发出高音调的蜂鸣声在151毫秒时发出低音调的蜂鸣声200毫秒时发出低音调的蜂鸣声250毫秒的高音调蜂鸣声有没有办法在Ruby或Python中做到这一点?我真的不在乎输出编码是什么(.wav、.mp3、.ogg等等),但我确实想创建一个输出文件。

  8. ruby-on-rails - Rails 3 中的多个路由文件 - 2

    Rails2.3可以选择随时使用RouteSet#add_configuration_file添加更多路由。是否可以在Rails3项目中做同样的事情? 最佳答案 在config/application.rb中:config.paths.config.routes在Rails3.2(也可能是Rails3.1)中,使用:config.paths["config/routes"] 关于ruby-on-rails-Rails3中的多个路由文件,我们在StackOverflow上找到一个类似的问题

  9. ruby-on-rails - 如何验证 update_all 是否实际在 Rails 中更新 - 2

    给定这段代码defcreate@upgrades=User.update_all(["role=?","upgraded"],:id=>params[:upgrade])redirect_toadmin_upgrades_path,:notice=>"Successfullyupgradeduser."end我如何在该操作中实际验证它们是否已保存或未重定向到适当的页面和消息? 最佳答案 在Rails3中,update_all不返回任何有意义的信息,除了已更新的记录数(这可能取决于您的DBMS是否返回该信息)。http://ar.ru

  10. ruby-on-rails - 'compass watch' 是如何工作的/它是如何与 rails 一起使用的 - 2

    我在我的项目目录中完成了compasscreate.和compassinitrails。几个问题:我已将我的.sass文件放在public/stylesheets中。这是放置它们的正确位置吗?当我运行compasswatch时,它不会自动编译这些.sass文件。我必须手动指定文件:compasswatchpublic/stylesheets/myfile.sass等。如何让它自动运行?文件ie.css、print.css和screen.css已放在stylesheets/compiled。如何在编译后不让它们重新出现的情况下删除它们?我自己编译的.sass文件编译成compiled/t

随机推荐